V4 generation வந்து ஒரு மாதம் கூட ஆகவில்லை, DeepSeek பேச்சை மாற்றிவிட்டது. செப்டம்பர் 10, 2026 அன்று DeepSeek-V4.1-Flash அறிமுகமானது – புத்தம் புதிய architecture குடும்பத்தின் மிகச்சிறிய மாடல், native vision, 1M-token context window, MIT license-இல் open weights உடன். வேண்டுமென்றே பரபரப்பூட்டும் கூற்று: மலிவான "Flash" மாடல் இப்போது agentic benchmark-களில் flagship V4 Pro-வை முந்துகிறது, சுமார் கால் விலையில்.

Launch post, pricing page, technical report evaluation tables அனைத்தையும் நான் படித்துவிட்டேன் – உங்களுக்காக இரண்டு விஷயங்கள் சேர்த்துள்ளேன்: முழு விலைக் கணக்கு (peak vs off-peak, cache hits, migration தேதிகள்) மற்றும் இன்றைய frontier மாடல்கள் – GPT-6 Astra, Claude Fable 5.1, GLM 5.3 – உடன் நேர்மையான benchmark ஒப்பீடு.

மூலக் குறிப்பு: கீழே உள்ள specs, விலைகள், scores அனைத்தும் DeepSeek-இன் அதிகாரப்பூர்வ அறிவிப்பு, அதிகாரப்பூர்வ API pricing page, model card evaluation tables (அனைத்தும் maximum reasoning effort-இல்) ஆகியவற்றிலிருந்து. இவை vendor-reported எண்கள், சுயாதீன VelsTech benchmark-கள் அல்ல – எங்கு இது முக்கியமோ அங்கு தெளிவாகக் குறிப்பிட்டுள்ளேன்.

சுருக்கமாக

V4.1 Flash என்றால் என்ன

V4.1-Flash என்பது DeepSeek-இன் V4.1 குடும்பத்தின் முதல், மிகச்சிறிய உறுப்பினர் – images, text படித்து text எழுதும் multimodal Mixture-of-Experts மாடல். 45T token-களில் scratch-இலிருந்து train செய்யப்பட்டது, thinking + non-thinking modes, cost-க்கும் accuracy-க்கும் இடையே trade-off செய்யும் dial (reasoning_effort 1–100), MIT license-இல் open-source – weights Hugging Face-இல் உள்ளன.

SpecDeepSeek-V4.1-FlashV4-Flash (முந்தைய தலைமுறை)V4-Pro
மொத்த parameters552B284B1.6T
Active / token8B prefill / 16B decode13B49B
ArchitectureCausal Encoder-Decoder, 40 layers (20+20)MoE + hybrid attentionMoE
Experts384 routed + 1 shared, 6 active
Context / max output1M / 384K1M / 384K1M / 384K
VisionNative (DeepSeek-ViT)தனி exp. variantஇல்லை
KV cache / token~890 bytes (FP4)~4x பெரியது
LicenseMIT (open weights)MITMIT weights / hosted API

இரண்டு engineering யோசனைகள் பெரும்பாலான வேலையைச் செய்கின்றன. Compressed Sparse Attention v2 key-value தரவை layer-களில் மீண்டும் compute செய்யாமல் share செய்கிறது, SWA Bounded Replay short-window attention-ஐ store செய்யாமல் தேவைக்கேற்ப reconstruct செய்கிறது – persistent cache V4-Flash-இன் HBM-இல் சுமார் 1/4, SSD-இல் 1/8, அசல் V1-இன் ~1/437 பங்கு. நாள் முழுதும் பெரிய context-களை மீண்டும் படிக்கும் agent workload-களில் பணம் இங்குதான் மிச்சமாகிறது.

விலை விவரம்: முழு கணக்கு

DeepSeek peak/off-peak pricing-ஐ வைத்து efficiency சேமிப்பை நேரடியாக rate card-இல் தந்துவிட்டது. புதிய விலைகள் செப்டம்பர் 10, 2026 முதல். Peak hours 01:00–04:00 மற்றும் 06:00–10:00 UTC, திங்கள்–வெள்ளி – மற்ற அனைத்தும் off-peak, சரிபாதி விலை.

1M token-க்குV4.1 Flash peakV4.1 Flash off-peakV4 Pro peak
Input, cache miss$0.30$0.15$1.32
Input, cache hit$0.006$0.003$0.044
Output$1.20$0.60$3.96

மூன்று விஷயங்கள் கவனிக்கத்தக்கவை. முதல், cache-hit தள்ளுபடி 50x: திரும்பத் திரும்ப வரும் system prompts, repo context, tool definitions $0.30-க்குப் பதில் $0.006/MTok. Context reuse செய்யும் agent-கள் – அதாவது கிட்டத்தட்ட அனைத்தும் – sticker price-ஐ விட மிக மலிவாகும். இரண்டு, V4.1-Flash peak-இல் V4 Pro-வை விட input-இல் ~4.4x, output-இல் ~3.3x மலிவு, off-peak-இல் இடைவெளி இரட்டிப்பு. மூன்று, migration தானாக நடக்கும்: model-ஐ deepseek-flash என set செய்யுங்கள்; பழைய deepseek-v4-flash, deepseek-v4-flash-vision-exp பெயர்கள் Flash rates-இல் V4.1-Flash-க்கு route ஆகும்.

Launch post-இல் ஒரு திருத்தம்: செப்டம்பர் 14, 04:00 UTC முதல் அனைத்து deepseek-v4-pro traffic-உம் Flash-க்கு மாறும் என்றார்கள். Pricing page இப்போது V4 Pro அந்தத் தேதிக்குப் பிறகும் மாறாத billing-உடன் தொடரும் என்கிறது – "Pro மூடப்படுகிறது" வரியை superseded எனக் கொள்ளுங்கள். Concurrency limits Flash-க்கு 2500 vs Pro-வுக்கு 500 – Flash-ஐ default ஆக்க இன்னொரு அமைதியான காரணம்.

Frontier உடன் benchmarks: Flash எங்கு வெல்கிறது, எங்கு இல்லை

கீழே உள்ள table DeepSeek-இன் சொந்த ஒப்பீடு (maximum reasoning effort, temperature 1.0 / top_p 0.95, code agent-களுக்கு DeepSeek Harness minimal mode). Rival columns: Opus-5.0, GPT-5.6 Sol, K3, GLM-5.3, V4-Pro, V4-Flash. "DeepSeek-இன் best foot forward" எனப் படியுங்கள் – harness, effort settings home team-க்கு சாதகம், சுயாதீன reruns பின்னர் வரும்.

BenchmarkV4.1 FlashTable-இல் best rivalமுடிவு
Terminal-Bench 2.1 (agentic terminal)90.6%89.1% (Opus-5.0)Flash முன்னிலை
DeepSWE v1.1 (software engineering)74.2%74.0% (Opus-5.0)Flash முன்னிலை
CyberGym88.1%84.5% (Sol / GLM-5.3)Flash முன்னிலை
HLE with tools63.9%63.6% (Opus-5.0)Flash முன்னிலை
AutomationBench (office workflows)54.8%50.3% (Opus-5.0)Flash தெளிவான முன்னிலை
Agent's Last Exam31.8%28.6% (Opus-5.0)Flash முன்னிலை
Codeforces rating34713348 (V4-Pro)Flash முன்னிலை
MathArena Apex65.6% (K3 உடன் tie)65.6% (K3)கூட்டு best
GPQA Diamond (reasoning)90.9%94.1% (GPT-5.6 Sol)Flagship-களுக்குப் பின்
HLE, tools இல்லாமல்36.8%56.3% (Opus-5.0)பின் – பெரிய இடைவெளி
Terminal-Bench 4.031.2%51.8% (Opus-5.0)புதிய harness-இல் பின்
SEC-Bench Pro62.8%74.3% (GPT-5.6 Sol)பின்
ExploitGym15.3%33.7% (GPT-5.6 Sol)பின்
ProgramBench / NL2Repo20.3% / 64.0%37.0% / 75.3% (Opus-5.0)பின்

Pattern சீரானது: "நீண்ட multi-step வேலை" bench-களில் V4.1-Flash ஆதிக்கம் – terminal agent-கள், repo-scale coding, automation, tool-using தேர்வுகள் – ஆனால் closed-book reasoning-இல் (tools இல்லா HLE), புத்தம் புதிய harness-களில் (Terminal-Bench 4.0, SEC-Bench) பெரிய closed மாடல்களுக்குப் பின். சொந்தக் குடும்பத்துக்கு எதிராகக் கதை தெளிவு: கிட்டத்தட்ட அனைத்து agentic row-களிலும் V4-Pro-வை வெல்கிறது, base-model coding-இலும் (HumanEval 79.4%, BigCodeBench 60.6%).

இந்த site-இன் புதிய flagship-களுடன் தொடர்பு என்ன? GPT-6 Astra (ARC-AGI-3 99.9%, OSWorld 2.0 72.6%, Terminal-Bench 4.0 57.9%), Claude Fable 5.1 (Terminal-Bench 4.0 55.8%, HLE with tools 65.0%) தங்கள் வலுவான harness-களில் இன்னும் முன்னிலை – ஆனால் அந்த எண்கள் வெவ்வேறு lab, harness, தேதிகளிலிருந்து, எனவே cross-vendor ஒப்பீட்டை directional ஆகக் கொள்ளுங்கள், controlled race ஆக அல்ல. நியாயமான சுருக்கம்: chat-இல் மட்டுமல்ல, agent வேலையில் flagship-களுடன் போட்டியிடும் முதல் open-weights மாடல் Flash.

Frontier விலைப் போர்: Flash vs அனைவரும்

Benchmarks பாதி மசாலாதான். அதே வேலை ஒரு மில்லியன் token-க்கு என்ன விலை (standard/peak rates; commit செய்யும் முன் ஒவ்வொரு vendor page-ஐயும் சரிபாருங்கள்):

ModelInput / 1MOutput / 1MCache read / 1M
DeepSeek V4.1 Flash (off-peak)$0.15$0.60$0.003
DeepSeek V4.1 Flash (peak)$0.30$1.20$0.006
DeepSeek V4 Pro (peak)$1.32$3.96$0.044
GLM-5.3$1.40$4.40$0.26
Gemini 2.5 Pro (≤200K)$1.25$10.00$0.125
GPT-5.6 Sol$4.00$20.00$0.40
GPT-6 Astra$10.00$50.00$1.00
Claude Fable 5.1$10.00$50.00$0.25

ஒரு realistic agent workload-க்குக் கணக்கிடுவோம் – 80% cache hits உடன் 200K input tokens, 50K output tokens கொண்ட பத்து request-கள், peak rates-இல் (அனைத்தும் 1M context-க்குள்). V4.1 Flash-இல் இது மொத்தம் சுமார் (0.4M × $0.30) + (1.6M × $0.006) + (0.5M × $1.20) ≈ $0.73. $1.00 cache reads உள்ள $10/$50 flagship-இல் அதே workload சுமார் $4.00 + $1.60 + $25.00 ≈ $30.60 – சுமார் 42x அதிகம். Off-peak-இல் Flash bill மீண்டும் பாதி – சுமார் $0.36. உங்கள் token mix-உடன் எங்கள் AI API cost calculator-இல் நீங்களே கணக்கிடுங்கள்; இந்த விலைகளில் self-hosting break-even நிறைய மாறும்.

Local AI பயன்படுத்துவோருக்கு இதன் அர்த்தம்

முதலில் நேர்மையான பதில்: 552B வீட்டில் ஓடாது. Safetensors release sparsely-accessed memory உட்பட ~763B parameters, third-party மதிப்பீடுகளின்படி full-precision serving-க்கு short context-இலேயே ~1,160 GB VRAM – சுமார் பதினெட்டு 80 GB datacenter card-கள். DeepSeek பெரிய scale-க்கு 2,000-GPU-plus-storage deployment-களைப் பற்றிப் பேசுகிறது. என் 12 GB RX 6800M இதை download செய்யாது; வீட்டில் உண்மையில் என்ன fit ஆகும் என்பதற்கு LLM-களுக்கு எவ்வளவு VRAM, local-LLM GPU guide பாருங்கள்.

நீங்கள் பயன்படுத்தக்கூடியது: weights MIT-licensed, எனவே FP8 originals உடன் community GGUF quant-கள், Ollama cloud entries (deepseek-v4.1-flash:cloud ஏற்கனவே உள்ளது), vLLM/SGLang support வேகமாக வரும். OpenCode, WorkBuddy (CodeBuddy உட்பட) ஏற்கனவே V4.1-Flash support-ஐ list செய்கின்றன – இந்த site வாசகர்களுக்கு முக்கியம், ஏனெனில் OpenCode-style harness-களில்தான் Flash-இன் benchmark பலம் (DeepSWE, Terminal-Bench, AutomationBench) வெளிப்படும். Local-AI நண்பர்களுக்கு practical move: private/sensitive வேலையை local 7B–35B மாடல்களில் வையுங்கள், heavy agentic வேலைகளை $10 API-க்குப் பதில் $0.15/MTok API-க்கு அனுப்புங்கள்.

முயற்சி செய்வது எப்படி

அனைத்தையும் மாற்றும் முன் எச்சரிக்கைகள்

Bottom line

V4.1-Flash இரண்டு எதிரெதிர் காரணங்களுக்காக ஒரே நேரத்தில் சுவாரஸ்யமான அரிய release. Architecturally, encoder-decoder + sparse-attention + replay தந்திரங்கள் 552B மாடலை 8B குறுகிய வழியில் squeeze செய்கின்றன – KV-cache சேமிப்பு (1/4 HBM, 1/8 SSD) உண்மையிலேயே புதியது. Economically, ஒவ்வொரு flagship-ஐயும் 30–60x undercut செய்கிறது, agent builder-களுக்கு முக்கிய bench-களில் சொந்த Pro-வையும் வெல்கிறது. உங்கள் workload பெரிய reused context-களுடன் நீண்ட tool-using session-கள் என்றால், இதுதான் இப்போதைய default to beat – open weights உட்பட. Closed-book Olympiad reasoning புதிய harness-இல் என்றால், flagship subscription-ஐ வைத்திருங்கள்.

Sources