Deepseek V4 Flash Vision Exp Local Deployment
Covered by
reddit.com
Timeline
- DeepSeek-V4-Flash-Vision-Exp runs on 10-12 RTX 3090s at 60-120 tok/s
A reproducible path to running a 285B MoE vision model with speculative decoding at 60-120 tok/s on consumer RTX 3090s could shift what counts as feasible for local, offline deployment of frontier-scale models.
1 source · thin-sourcing, unconfirmed
All sources (1)
- r/LocalLLaMA2026-09-09
Related subjects
- Cosmos3 Int4 Quant Local ReleaseLocal AI Scene
- Mentria Browser Inference EngineLocal AI Scene
- Minnow Llada2 2 Inference ServerLocal AI Scene