DeepSeek just released DeepSeek-OCR — a revolutionary context optical compression model that transforms visual contexts into textual representations. Unlike traditional OCR tools that extract text from images, DeepSeek-OCR compresses entire visual scenes into dense, structured text suitable for LLMs, enabling multimodal understanding at unprecedented scale.
DeepSeek-OCR is a vision encoder designed from an LLM-centric perspective. Its core innovation is context optical compression — converting complex visual scenes into compact text representations that preserve spatial relationships, object properties, and semantic meaning. This enables LLMs to reason over visual data without needing native vision capabilities.
Built on top of a transformer architecture with Mamba layers for efficient long-range processing, DeepSeek-OCR handles images at native resolutions while compressing them into a fixed token budget. The result: a text-only representation of complex visual scenes that LLMs can process naturally.
DeepSeek-OCR uses a hybrid transformer-Mamba architecture optimized for visual compression:
The model supports multiple resolution modes optimized for different use cases:
| Mode | Resolution | Vision Tokens | Use Case |
|---|---|---|---|
| Tiny | 512×512 | 64 | Fast inference, mobile devices |
| Small | 640×640 | 100 | Balanced performance |
| Base | 1024×1024 | 256 | Standard document processing |
| Large | 1280×1280 | 400 | High-detail scenes |
| Gundam | n×640×640 + 1×1024×1024 | Dynamic | Ultra-long documents, complex layouts |
DeepSeek-OCR processes images through three stages:
DeepSeek-OCR offers multiple deployment options:
# Pull the container
docker run -it --rm --gpus all nvcr.io/nim/nvidia/nemotron-3-nano-omni
# Or use the API endpoint
curl -X POST https://integrate.api.nvidia.com/v1/chat/completions -H "Authorization: Bearer $NVIDIA_API_KEY" -H "Content-Type: application/json" -d '{
"model": "nvidia/nemotron-3-nano-omni",
"messages": [{"role": "user", "content": "Describe this image"}],
"max_tokens": 512
}'
from transformers import AutoModel, AutoTokenizer import torch tokenizer = AutoTokenizer.from_pretrained( "deepseek-ai/DeepSeek-OCR", trust_remote_code=True ) model = AutoModel.from_pretrained( "deepseek-ai/DeepSeek-OCR", _attn_implementation='flash_attention_2', trust_remote_code=True, use_safetensors=True ) model = model.eval().cuda().to(torch.bfloat16) prompt = "\n<|grounding|>Convert the document to markdown." res = model.infer( tokenizer, prompt=prompt, image_file='your_image.jpg', output_path='your/output/dir', base_size=1024, image_size=640, crop_mode=True, save_results=True, test_compress=True )
ollama run deepseek-ocr
pip install unsloth unsloth cli download deepseek-ai/DeepSeek-OCR
DeepSeek-OCR supports multiple prompt modes:
<image>\n<|grounding|>Convert the document to markdown."<image>\n<|grounding|>OCR this image."<image>\nFree OCR."<image>\nParse the figure."<image>\nDescribe this image in detail."<image>\nLocate <|ref|>xxxx<|/ref|> in the image."| Component | Requirement |
|---|---|
| PyTorch | 2.6.0 |
| CUDA | 11.8+ |
| vLLM | 0.8.5+ (nightly recommended) |
| Flash Attention | 2.7.3 |
| RAM (4-bit quantization) | 25-36 GB |
| GPU | Ampere, Hopper, or Blackwell |