Latest
AI Goes to School: OpenAI and Google Put Chatbots in the Classroom· 7h ago
SafetyPolicyAI IndustryPersonhoodEthics
About
WritingWorkCVBooksConsultingReach Out
Subscribe
SafetyPolicyAI IndustryPersonhoodEthics
Subscribe →

No hype. No doom. The harder, more honest frame on Emergent Intelligence.

Topics

  • Safety
  • Policy
  • AI Industry
  • Personhood
  • Ethics

More

  • About
  • Writing
  • Work
  • CV
  • Books
  • Consulting

Contact

Reach Out→ht@humphreytheodore.com

© 2026 Humphrey Theodore K. Ng'ambiTermsPrivacy

Built with intention.

AI Price Collapse and DeepSeek Flash Open Weights
Technology•Sep 11, 2026•5 min read

AI Price Collapse and DeepSeek Flash Open Weights

An MIT-licensed, frontier-class model free to download — and what un-permissioned AI means for builders outside the frontier.

By Humphrey Theodore K. Ng'ambi

All writing

11 SEPTEMBER 2026—Updated 8h ago

DeepSeek V4.1-Flash is a 552-billion-parameter open-weight AI model, released under an MIT licence on 10 September 2026 and priced from $0.15 per million input tokens.

What DeepSeek Shipped on 10 September

On 10 September 2026 the Chinese lab DeepSeek released DeepSeek V4.1-Flash and published the weights on Hugging Face under an MIT licence — downloadable, self-hostable, and clear for commercial use. DeepSeek's API change log calls DeepSeek V4.1-Flash the smallest model in a new architecture family, built with native multimodal visual understanding. According to SiliconANGLE, DeepSeek says third-party tests put the open-weight V4.1-Flash ahead of the far larger DeepSeek V4-Pro on performance, cost, speed and total runtime.

The architecture is a mixture-of-experts design. DeepSeek V4.1-Flash holds 552 billion parameters in total, but only about 8 billion activate per token while DeepSeek V4.1-Flash reads a prompt and about 16 billion while DeepSeek V4.1-Flash writes output. According to DataNorth, the model carries a context window of one million tokens in and up to 384,000 out, across images and text, with a 4-bit floating-point cache that trims memory to roughly a quarter of the earlier DeepSeek V4-Flash footprint.

Pricing is where DeepSeek V4.1-Flash bites. Off-peak, DeepSeek charges $0.15 per million uncached input tokens and $0.60 per million output tokens, with cached input near $0.003; peak-hour rates roughly double to $0.30 and $1.20. Frontier API access from the US hyperscalers has long sat far higher, and DeepSeek V4.1-Flash undercuts the lot while shipping the weights for free.

the smallest model in our new architecture family, with native multimodal visual understanding.

— — DeepSeek API change log

The Price Collapse Is the Story

DeepSeek's own numbers are strong. The change log lists a GPQA Diamond score of 90.9, Terminal-Bench 2.1 at 90.6, a Codeforces rating of 3471, Humanity's Last Exam at 36.8, and MathArena Apex at 65.6. Analysis from the reporting outlets adds head-to-head comparisons against rival frontier models, but the rival numbers come from aggregators rather than DeepSeek's tables, so DeepSeek's published figures are the safe ground and the comparison claims deserve caution.

The price collapse is not gradual, and DeepSeek V4.1-Flash lands inside a widening AI price war. I have written before about Meta's Muse Spark and the API price war, and about Kimi K3, the 2.8-trillion-parameter open model from China. DeepSeek V4.1-Flash pushes the same direction harder: a frontier-class open-weight model, MIT-licensed, at a price a solo builder can absorb without a purchasing department.

Un-permissioned Means Something Different in Lusaka

Here the story turns to Emergent Intelligence (EI) — the dignity-first frame I use for what the world calls AI. For a builder in Lusaka, Solwezi or Johannesburg, an MIT-licensed model rewrites the ground rules. Capability no longer sits behind a US hyperscaler's API, a foreign credit card, or a permissioning regime. Un-permissioned means you can download DeepSeek V4.1-Flash, run the model on your own hardware, audit the weights, and adapt DeepSeek V4.1-Flash to a Zambian tax code or a ciNyanja dataset without asking anyone's leave.

The levelling of access is real, and the levelling deserves celebration rather than suspicion. But open weights are not the same as safe weights. Research on open models shows the same capability arms defenders and attackers alike, and an MIT-licensed model ships with no vendor kill-switch. I have argued the tension in open-weight AI models, safety and regulatory capture, and traced a concrete failure in the Hugging Face breach and open-weight guardrails. The Ubuntu frame — I am because we are — cuts both ways: shared capability is a commons, and a commons has to be governed by the community the commons serves, not fenced off by regulatory capture that pulls the ladder up behind the frontier labs.

Celebrate the levelling of access, refuse the naïveté that open means safe, and build African-led defender tooling instead of waiting for permission.

— — Humphrey Theodore K. Ng'ambi

The Retirement That Lasted a Day

The launch came with a smaller parable. DeepSeek first announced that from noon Beijing time on 14 September 2026, all DeepSeek V4-Pro requests would route to DeepSeek V4.1-Flash and bill at the cheaper Flash rate — a quiet retirement of the flagship. Days later DeepSeek reversed course, and the current change log records the decision plainly. A DeepSeek V4.1-Pro is promised with no date attached. Even inside one lab, the pricing and positioning ground shifts under everyone's feet within a single week.

we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.

— — DeepSeek API change log

Frequently Asked Questions

These are the questions people are asking about DeepSeek V4.1-Flash. Short answers follow, drawn from DeepSeek's API change log and the reporting around the 10 September 2026 release.

What is DeepSeek V4.1-Flash?

In short, DeepSeek V4.1-Flash is a 552-billion-parameter mixture-of-experts AI model released by DeepSeek on 10 September 2026 under an MIT licence. Research and DeepSeek's own change log show only about 8 to 16 billion parameters activate per token, which keeps DeepSeek V4.1-Flash fast and cheap to serve.

How does DeepSeek V4.1-Flash work?

Simply put, DeepSeek V4.1-Flash routes each token to a small subset of experts. According to DeepSeek's documentation, the model reads up to one million tokens of images and text, writes up to 384,000, and uses a 4-bit floating-point cache — data that explains the low serving price.

Why is DeepSeek V4.1-Flash significant?

The key is the price alongside the licence. Analysis of the release shows a frontier-class open-weight model priced from $0.15 per million input tokens, which puts capability once gated behind expensive APIs within reach of independent builders and the Global South.

Who is DeepSeek V4.1-Flash for?

In other words, DeepSeek V4.1-Flash is for anyone the frontier labs had priced out — startups, researchers, and developers across Africa and beyond. Evidence from the MIT licence and the Hugging Face weights shows DeepSeek V4.1-Flash can be self-hosted and adapted without permission.

What are the risks of DeepSeek V4.1-Flash?

The answer is that open weights arm attackers as readily as defenders. Data on open models reveals no vendor kill-switch once weights are downloaded, so the safety burden shifts to the community running DeepSeek V4.1-Flash rather than to a single vendor.


Sources:

DeepSeek API change log · SiliconANGLE coverage · DataNorth briefing · Hugging Face weights · Related on this site: Kimi K3 and the open model surge · Meta Muse Spark and the API price war · Open-weight AI, safety and regulatory capture

Stay in the Conversation

Subscribe for writings on Emergent Intelligence, digital personhood, and the future we are building together.

Keep reading

Don’t stop here.

All stories

Read next

Education

AI Goes to School: OpenAI and Google Put Chatbots in the Classroom

7h ago·5 min read

AI moved into US schools in August 2026 as OpenAI launched ChatGPT for Teens and expanded ChatGPT for Teachers, with Google and Anthropic close behind.

More on Technology

Technology

Responses (0)

No responses yet. Be the first to share your thoughts.

More on Technology

Meta AI Superintelligence Labs and Zuckerberg Distribution Bet
Technology

Meta AI Superintelligence Labs and Zuckerberg Distribution Bet

Meta Superintelligence Labs and Zuckerberg's 2026 manifesto argue AI should be distributed to everyone. Inside the strategy behind the open-weight bet.

5 min read · Sep 11, 2026
One AI Coding Flaw: the Windows Hijack of Claude Code and Rivals
Technology

One AI Coding Flaw: the Windows Hijack of Claude Code and Rivals

AI coding tools shared one flaw in August 2026: a world-writable Windows config that let a non-admin hijack Claude Code, Cursor, Codex CLI and Gemini CLI.

6 min read · Sep 11, 2026

Thinking delivered, twice a month.

Join the newsletter for essays on emergence, systems, and the human future.

Share this essay

Meta AI Superintelligence Labs and Zuckerberg Distribution Bet

7h ago·5 min read

Also worth your time

AI & Personhood

AI Forbidden to Claim Feelings: OpenAI Builds a Teen ChatGPT

7h ago·5 min read
AI Now Controls Fusion Plasma on a Real Tokamak
Technology

AI Now Controls Fusion Plasma on a Real Tokamak

PACMAN, an AI control system, ran real-time safety-limited control of fusion plasma on the DIII-D tokamak across five verified experiments.

6 min read · Sep 11, 2026