Overview
Restricted model access dominates the release news, but the day's stronger evidence concerns how AI is used and measured. OpenAI and Anthropic telemetry quantify delegation, a 67-model study locates a co-failure ceiling, and TreasuryBench's open artifacts expose both consequential errors and a vendor's conflict in grading its own product.
1. GPT-5.6 Sol enters a limited preview at $5 input and $30 output
OpenAI opened GPT-5.6 Sol, Terra, and Luna to a small group of trusted partners whose participation was shared with the US government. Per million tokens, Sol costs $5 input and $30 output, Terra $2.50 and $15, and Luna $1 and $6; broader availability is planned within weeks.
Sol adds a max reasoning setting and an ultra mode that coordinates subagents. OpenAI classifies all three models as High capability for cybersecurity and biological or chemical risk, below its Critical cyber threshold, and reports spending more than 700,000 A100-equivalent GPU hours on automated jailbreak discovery.
Sources: OpenAI · OpenAI Deployment Safety
2. Mythos reportedly reaches selected US organizations after access restrictions
Anthropic's Mythos became available to selected trusted US organizations with government permission, according to a Reuters report citing Semafor. The report followed restrictions that had removed the model from broader access, turning eligibility into a condition of deployment even when the model checkpoint remained unchanged.
Together with GPT-5.6's government-informed partner list, the Mythos release establishes a selective channel outside ordinary product tiers. The public report omits selection criteria, monitoring terms, incident procedures, and expansion conditions, leaving prospective customers unable to distinguish a temporary evaluation cohort from an emerging licensing pattern.
Sources: Reuters
3. Claude Code raises measured autonomy by 0.37 points across output types
Anthropic's Economic Index rated AI autonomy on a five-point scale and found Claude Code averaged 0.37 points above chat and Cowork across output types. Claude Code used Opus in 54% of sampled sessions versus 10% elsewhere, but a same-model comparison on Sonnet retained a 0.26-point gap.
The report also linked usage data to about 9,700 survey respondents: more than 35% expected AI to perform most of their work tasks within 12 months, while 10% rated losing their own job likely or very likely. Computer and mathematical workers comprised roughly 30% of respondents but 4% of US employment, sharply limiting population-level inference.
Sources: Anthropic
4. Codex users increasingly delegate tasks estimated above eight human hours
Weekly active Codex users grew more than fivefold from January 1 to June 1. During the 28 days ending June 11, 17.3% of active organizational users tried Codex, yet they generated 63.3% of their combined Codex and ChatGPT output tokens; the corresponding individual-account shares were 0.7% and 16.5%.
In a 0.1% sample of opted-in individual accounts, the share of users submitting at least one task classified above eight hours of human work rose from 2.1% in December to 25.6% in May. The duration comes from a model estimator, and OpenAI's nearly frictionless internal adoption is explicitly presented as an upper-bound setting, not a representative workplace.
Sources: OpenAI
5. Shared failures set a 94.8% ceiling for one 67-model math pool
A preprint tests 67 AI models from 21 provider families and measures an all-models-wrong rate of 5.2% on open-ended MATH-500, versus 2.3% predicted by its fitted Gaussian model. That co-failure rate caps any router, vote, or cascade selecting a member answer at 94.8%; execution-graded code produced a 7.9% co-failure rate.
On a separate 15-model mix, a learned text router scored 90.6% against 90.1% for the best single model, realizing 9% of the oracle gain with a confidence interval spanning zero; an LLM router chose the single-best model for every query. The single-author preprint's limited all-wrong counts make its open artifacts and replication more informative than the headline alone.
Sources: arXiv
6. Treasury's own benchmark ranks Treasury first on personal-finance advice
TreasuryBench evaluates 81 AI personal-finance tasks across three synthetic households and 12 domains. Its published run scores Treasury at 85.5, a full-context ChatGPT chat-latest baseline at 79.6, Origin at 71.0, and Monarch at 52.1; the factual checks flag one dangerous Treasury error and 12 for ChatGPT.
Treasury Technologies authored the benchmark and the product that ranks first, creating a direct self-evaluation conflict despite the use of an AI judge. The repository discloses self-authorship and publishes captures and judgments, but the creator still selected the personas, planted facts, tasks, scoring rules, and product run; Origin also lost balance imports on 16 tasks.