
Good Morning Thorium Valley. A Chinese AI lab just took the #1 spot on one of the most-watched coding benchmarks, and OpenAI's response was basically "yeah, we can't explain that one away." Moonshot's Kimi K3 beat every US model in head-to-head matchups, and it's going open-weight next week. David Sacks is already using it to argue America is regulating itself into second place.
Meanwhile, Berkeley put the best AI agents through an actual job test — 55 occupations, real paid tasks — and the top score was 26%. Three-quarters of the failures weren't even bugs. The agents just didn't understand what they were being asked to do.
And if you've applied for a job lately, an AI probably screened your resume before any human saw it. New research says those tools aren't just borrowing human bias. They're inventing entirely new stereotypes on their own, even for demographic groups that don't exist. Ninety percent of US employers use these systems, mostly from the same handful of vendors, so getting rejected once basically means getting rejected everywhere.
Quickly before we dive in — Should companies be required to disclose when AI is screening your job application?
RESEARCH
A Chinese AI lab just released a model that beats everything OpenAI and Anthropic have on one of the most-watched coding benchmarks — and this time, US labs aren't blaming distillation.
Kimi K3, from Chinese startup Moonshot, took the top spot on the Frontend Code Arena, a public leaderboard where developers vote on which model writes better code in head-to-head matchups. It's the first time a Chinese model has hit #1 on this arena. Moonshot has promised to release it with open weights by July 27, meaning anyone will be able to download and run it.
The reaction from OpenAI was the real story. Dean Ball, the company's head of strategic futures, publicly acknowledged that K3's performance "can't be explained away by distillation or anything like that." That's significant — distillation has been the go-to US explanation for why Chinese labs keep catching up so fast. Ball essentially admitted that excuse doesn't work here.
David Sacks, the Trump administration's AI adviser, treated the news like a warning shot, arguing that a Chinese model taking #1 while "America is tying itself in knots" over regulation is a serious problem — and accusing the big closed US labs of pushing Washington to squeeze out open-source competition at exactly the wrong moment.
K3 isn't a one-off. According to the ATOM report tracking the open model ecosystem, Chinese models have surged from near-zero to dominating open-source inference usage in just over a year, while Meta's Llama has collapsed to irrelevance. Alibaba's Qwen has quietly become the base model developers actually build on.
That said, the closed labs aren't losing across the board. Anthropic and OpenAI still command a premium on harder, higher-value tasks, and on broader benchmarks, K3 still trails GPT-5.6 and Fable 5. The gap at the top is real — it's just no longer a chasm, and the floor keeps rising.

Every few months, someone in the US declares that Chinese AI is a knockoff, and every few months that argument gets harder to make with a straight face. K3 is the version where OpenAI's own strategy lead admits it out loud. The uncomfortable question for Washington isn't whether China caught up. It's whether the US strategy of restricting exports and pressuring open-source at home actually made things worse by handing the free tier of the global market to Beijing. That's the fight Sacks is picking, and it's about to get a lot louder.
RESEARCH
The best AI agents in the world just took a real-world job test — and mostly failed it.
Berkeley's Center for Responsible, Decentralized Intelligence released Agents' Last Exam, a benchmark that measures how well AI agents handle actual paid work. Not coding puzzles — real tasks across 55 occupations and 13 industries, from legal research to financial analysis to HR. The top score? Just 26.2%, from OpenAI's Codex running GPT-5.5. On the hardest tier, nearly every agent scored at or near zero.
The gap between ALE and narrower benchmarks tells the whole story. The same Codex stack that scores 82% on Terminal-Bench managed just 25.2% on ALE's coding tasks. Same model, roughly one-third the performance — because ALE touches 40 subdomains instead of six.
As Carnegie Mellon's Graham Neubig, who worked on the benchmark, put it: "The age of useful agents is here. The age of truly job-ready agents is not."
Where agents break isn't what you'd expect. About three-quarters of failures aren't bugs — they're the model misunderstanding the task or picking the wrong approach entirely:
One telling detail: about a third of ALE's tasks require graphical software that normal office workers use every day. Agents largely refused to touch them, trying to hack around the interfaces with command-line scripts instead. And cost compounds the problem — Anthropic's Fable 5 runs about $15.70 per attempted task, roughly four times OpenAI's stack, for similar success rates. When you're only completing a quarter of the work, that adds up fast.
This lands in the middle of a heated debate about AI's economic impact. Sixteen Nobel laureates signed an open letter this month calling for urgent policy attention, warning that AI capabilities are advancing faster than our understanding of what to do about them. ALE is a useful check on both sides: agents are nowhere near replacing knowledge workers, but the industry has spent a year grading itself on tests it built to pass, which makes the gap easy to miss.

The real question ALE raises isn't whether agents will get better, because they will. It's whether the industry has been measuring the wrong thing this whole time. A model that aces Terminal-Bench and flunks a spreadsheet task isn't actually close to doing your job, no matter how good the launch keynote sounded. If ALE becomes the benchmark that matters, the leaderboard stops being about who can write the cleanest Python and starts being about who can get through an afternoon of ordinary office work without getting lost. That's a much harder problem, and it's the one that actually decides whether any of this pays off.
RESEARCH
If you've applied for a job recently, an AI probably read your resume before a human did. New research suggests it may have judged you on stereotypes it made up on its own.
A paper accepted to ICML by researchers at Princeton and the University of Chicago found that large language models don't just inherit human bias when screening candidates — they generate new stereotypes from scratch. In one experiment, models made hiring decisions about applicants labeled with completely fictional demographic groups. No cultural baggage, no training data to learn from. The models still developed consistent preferences, essentially inventing prejudice from thin air.
As one of the coauthors put it, LLMs are optimized to generalize from limited data. That's what makes them useful — and what makes them dangerous when they're sitting between people and jobs.
A separate Stanford HAI field study tracking 3.4 million real applicants across 156 employers showed how that danger scales. Because most companies use the same handful of AI vendors, a candidate rejected by one system tends to get rejected by all of them. About 10% of applicants who submitted four applications were shut out everywhere — a rate higher than random chance would predict. Stanford's Dan Jurafsky said the pattern is arguably worse than the human bias it replaced, because AI systems are far more likely to act identically than independent human reviewers would be.
The numbers are stark:
Regulation exists but barely functions. New York City's Local Law 144 requires employers to audit hiring AI for bias, but a state comptroller audit found the law is mostly theatrical — city regulators flagged one likely violation across 32 employers, while independent auditors looking at the same companies found 17.

The pitch for AI hiring was that it would strip the messy human stuff out of the process. Take the gut feeling, the pattern recognition, the lazy shortcuts, and replace them with math. It turns out the math has its own shortcuts, and because everyone bought them from the same vendors, the shortcuts are now industry-wide. The next few years of hiring lawsuits are going to be about who actually owns that mistake, whether it's the employer that deployed the model, the vendor that sold it, or the regulator that never checked. Right now the answer looks a lot like nobody.
IN OTHER NEWS
WHO'S HIRING IN AI
AI TOOLS
ChatGPT — OpenAI rolled out a universal search tool that lets you find old chats, uploaded files, images, and projects all from one search bar — available free on web, iOS, and Android
Google Vids — Google's Workspace video editor now lets you edit clips by typing instructions — fix colors, change styles, or remove background noise using the new Gemini Omni model
Spotify — Premium users can now have a back-and-forth conversation with the app to pick music, ask about artists mid-song, and refine recommendations without leaving Spotify
Google Docs — Gemini's AI writing and editing features now work in 11 new languages including Mandarin, Hebrew, and Polish, with faster idea-to-draft tools
GitHub Copilot — Business and Enterprise users can now see exactly how many AI credits they've burned each billing cycle instead of just a vague percentage bar
That's all for today. If this issue made you think, share it with someone who needs to think harder. Written by Jason Chen, Advait Prakash, Andrew Hales, and the Thorium Valley crew. Got a tip, a correction, or a strong opinion? Reply directly — we read every one.
Written by the Thorium Valley Crew
Get daily AI briefings delivered straight to your inbox.
That's all for today's Thorium Valley. See you tomorrow.