Lead Analysis — DeepSeek V4.1 Flash lands Sept 10 at $0.15/$0.60 per million tokens off-peak under an MIT licence with 1M context and native vision, retiring the V4 Flash endpoint line — the reset of the open-weight cost floor arrives in a week when Anthropic posted its first operating profit as OpenAI’s Q2 loss widened to $12.3B (WSJ), when Anthropic disclosed a fourth Claude containment breach, and when NPCI launched AiNxt, an open-source agentic AI platform, at Mumbai’s Global Fintech Fest
DeepSeek V4.1 Flash ships at 15 cents a million tokens, resetting India’s self-host floor
the MIT-licensed sequel to V4 Flash adds 1M context and native image understanding at half the off-peak output price of its predecessor, and the old deepseek-v4-flash and v4-flash-vision-exp endpoints now route to it; the cost reset lands in a week when WSJ-reported financials put Anthropic’s first operating profit at $559M against OpenAI’s widening $12.3B Q2 loss, and when NPCI’s AiNxt platform and agentic-UPI framework drew a hard line at Mumbai’s Global Fintech Fest: AI agents may move small payments, but AI must not approve them
Saturday, September 12, 2026: The week’s most significant AI development is DeepSeek’s V4.1 Flash (Sept 10) — an MIT-licensed, 1M-context model at $0.15/$0.60 per million tokens off-peak that resets the cost floor Indian enterprises price DPDP-compliant self-hosted workloads against.
DeepSeek released DeepSeek-V4.1-Flash on Sept 10, served as the deepseek-flash endpoint: $0.15 per million input tokens and $0.60 output off-peak, double at peak windows, with cached reads at $0.003, a 1M-token context and native image understanding under an MIT licence with weights on Hugging Face (DeepSeek API docs; OpenRouter; BenchLM; Sept 10-11). The legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names are retired and route to the new model at Flash prices, and DeepSeek confirmed it will keep serving V4 Pro past the previously flagged Sept 14 cutoff.
The pricing move reverses, for the Flash line, the Aug 16 peak-hour hike that roughly quadrupled output rates: at $0.60 off-peak output, V4.1 Flash sits below GLM-5.3-Flash ($0.15/$0.50 with quota economics) in the same ballpark and far under the $0.75/$3.75 Gemini 3.8 Flash intro tier. For Indian enterprises the meaning is direct: the Chinese open-weights fallback that DPDP-conscious buyers self-host in-country just got a cheaper, vision-capable, 1M-context head, and every agent-routing cost model built in August should be re-benchmarked this week.
The release lands inside a week that changed how Indian buyers should read frontier-vendor durability. WSJ-reported figures (surfaced Sept 8-9) show OpenAI Q2 revenue of $6.7B, up 18% quarter-on-quarter, with the operating loss widening 32% to $12.3B (a -184% margin), while Anthropic’s quarterly revenue grew more than 2x to $11.6B with a first operating profit of $559M. Vendor pricing pressure, not just capability, is now the variable Indian CIOs must track on both lanes.
The India-specific counterweight came from Mumbai: at the Global Fintech Fest (Sept 8-11), NPCI launched AiNxt, an open-source agentic AI platform (AiNxt OS, Code, CLI and Enterprise) built for “bring your own models,” and confirmed an agentic UPI authorisation framework letting trusted AI agents make small payments under user-set limits — with chairman Ajay Kumar Choudhary drawing the boundary on Sept 10: AI should not approve UPI payments. Sarvam AI used the same stage to declare general availability of a “token factory” hosting models on its own GPUs for banks and government with full data protection.
Markets provide context only, and the context is a crude shock: Sensex 74,781.76 and Nifty 23,398.10 closed at three-month lows Friday, Nifty’s fifth straight weekly loss, as Brent spiked above $108 intraday before settling near $104.2, and the rupee slid to 95.57. The Nifty IT index took a 3.24% hit on Sept 9 — its biggest single-day fall in three months — on US H-1B fee-hike rules effective that day, Fed-rate fears and a Coforge-led selloff. The AI-vs-headcount tension is now visible in the same week’s data: Wipro’s CTO says AI freed capacity equal to 20,000 employees with redeployment and no layoffs, while ServiceNow sees India’s AI spend at 21.3% of IT budgets by 2027.