Overview
OpenAI's reported expenditure illustrates the capital behind frontier research, but today's stronger evidence comes from task construction and physical or simulated outcomes. LifeSciBench, a 10,080-reaction chemistry program, and a 100-scenario care study expose sharply different capability boundaries.
1. OpenAI's reported spending reaches $34 billion for the year
OpenAI spent about $34 billion last year, the Financial Times reported, citing financial documents. The figure comes from reported internal records, not an audited public filing, but it illustrates the scale of compute, research, product, and infrastructure commitments behind frontier-model development.
The reported total covers compute, research, products, and infrastructure without the detail of an audited public filing. It puts provider continuity and future pricing against a spending base that must be financed before successive training runs and serving commitments generate returns.
Sources: Financial Times
2. LifeSciBench turns research work into expert-reviewed model tasks
LifeSciBench contains 750 tasks, 1,062 artifacts, and 19,020 rubric criteria assembled by 173 scientists and reviewed by 453 independent experts. Seventy-nine percent of tasks contain multiple reasoning or decision steps, and 53% include at least one artifact.
GPT-Rosalind passed 36.1% overall, including 71.1% on a nine-task scientific-communication subset, but only 14.8% on numeric tasks. Its pass rate fell from 45.1% on text-only work to 28.1% when artifacts or URLs entered the task.
Sources: OpenAI
3. OpenAI reports GPT-5.4 improved yields in chemistry experiments
GPT-5.4 and Molecule.one's Maria system proposed TEMPO for a difficult primary-sulfonamide Chan-Lam coupling, then ran 10,080 reactions across two experimental cycles. Mean yield increased from 16.6% to 25.2%, while reactions above 30% yield rose from 15.6% to 37.5%.
Human chemists reproduced improved yields in 11 of 14 substrate pairs at bench scale, including greater than twofold gains in eight. Humans selected proposals, corrected plans, handled laboratory work, and validated the output, so the experiment measures a supervised research loop within one reaction family.
Sources: OpenAI
4. AMIE outperforms primary-care physicians on simulated longitudinal management
Nature reports a randomized, blinded virtual examination comparing the AMIE clinical AI system with 21 primary-care physicians across 100 scenarios, three visits per scenario, and five specialties. Specialist raters scored AMIE's plans as appropriate in 95%-98% of cases across visits, versus 72%-81% for physicians.
The text-only simulations used constructed cases, compressed visits into one or two days, and supplied UK-oriented guidelines to physicians based in Canada and India. On the separate RxQA medication test, peak open-book accuracy remained below 75% for both AMIE and physicians.