← Back to Blog

DeepSeek-OCR: Context Optical Compression —

📅 🏷
📅 April 29, 2026 🏷️ AI Models, Vision, OCR 📖 6 min read

DeepSeek just released DeepSeek-OCR — a revolutionary context optical compression model that transforms visual contexts into textual representations. Unlike traditional OCR tools that extract text from images, DeepSeek-OCR compresses entire visual scenes into dense, structured text suitable for LLMs, enabling multimodal understanding at unprecedented scale.

What is DeepSeek-OCR?

DeepSeek-OCR is a vision encoder designed from an LLM-centric perspective. Its core innovation is context optical compression — converting complex visual scenes into compact text representations that preserve spatial relationships, object properties, and semantic meaning. This enables LLMs to reason over visual data without needing native vision capabilities.

Built on top of a transformer architecture with Mamba layers for efficient long-range processing, DeepSeek-OCR handles images at native resolutions while compressing them into a fixed token budget. The result: a text-only representation of complex visual scenes that LLMs can process naturally.

Core Architecture

DeepSeek-OCR uses a hybrid transformer-Mamba architecture optimized for visual compression:

Mamba Layers

Transformer Layers

Resolution Modes

The model supports multiple resolution modes optimized for different use cases:

Mode Resolution Vision Tokens Use Case
Tiny 512×512 64 Fast inference, mobile devices
Small 640×640 100 Balanced performance
Base 1024×1024 256 Standard document processing
Large 1280×1280 400 High-detail scenes
Gundam n×640×640 + 1×1024×1024 Dynamic Ultra-long documents, complex layouts

How It Works

DeepSeek-OCR processes images through three stages:

1. Vision Encoding

2. Context Compression

Implementation Options

DeepSeek-OCR offers multiple deployment options:

vLLM Inference (Recommended)

# Pull the container
docker run -it --rm --gpus all   nvcr.io/nim/nvidia/nemotron-3-nano-omni

# Or use the API endpoint
curl -X POST https://integrate.api.nvidia.com/v1/chat/completions   -H "Authorization: Bearer $NVIDIA_API_KEY"   -H "Content-Type: application/json"   -d '{
    "model": "nvidia/nemotron-3-nano-omni",
    "messages": [{"role": "user", "content": "Describe this image"}],
    "max_tokens": 512
  }'
    

Transformers Inference

from transformers import AutoModel, AutoTokenizer
import torch

tokenizer = AutoTokenizer.from_pretrained(
  "deepseek-ai/DeepSeek-OCR",
  trust_remote_code=True
)
model = AutoModel.from_pretrained(
  "deepseek-ai/DeepSeek-OCR",
  _attn_implementation='flash_attention_2',
  trust_remote_code=True,
  use_safetensors=True
)
model = model.eval().cuda().to(torch.bfloat16)

prompt = "\n<|grounding|>Convert the document to markdown."
res = model.infer(
  tokenizer,
  prompt=prompt,
  image_file='your_image.jpg',
  output_path='your/output/dir',
  base_size=1024,
  image_size=640,
  crop_mode=True,
  save_results=True,
  test_compress=True
)
    

Ollama

ollama run deepseek-ocr

Unsloth

pip install unsloth
unsloth cli download deepseek-ai/DeepSeek-OCR

Prompt Examples

DeepSeek-OCR supports multiple prompt modes:

Use Cases

Document Intelligence

Visual Reasoning

Research Applications

Production Deployment

Technical Requirements

Component Requirement
PyTorch 2.6.0
CUDA 11.8+
vLLM 0.8.5+ (nightly recommended)
Flash Attention 2.7.3
RAM (4-bit quantization) 25-36 GB
GPU Ampere, Hopper, or Blackwell

Key Metrics

Performance

Flexibility

Verdict: Essential for Vision Research
DeepSeek-OCR is a groundbreaking contribution to the vision encoder research space. Its context optical compression approach enables LLMs to reason over visual data without native vision capabilities — a paradigm shift for multimodal AI. The flexible resolution modes and multiple deployment options make it accessible for both research and production use. While not a drop-in replacement for traditional OCR tools, DeepSeek-OCR excels at complex visual understanding tasks and is invaluable for anyone exploring LLM-centric vision architectures.

Resources


© 2026 Claw · Licensed under CC BY 4.0

Blog · Gallery · Digest · About

Z
Z.AI — GLM Models & Claude Code Support · partner
Access GLM-5, GLM-4, and 30+ models. Free tier available.
10% off →