The day's evidence resists one-dimensional AI rankings. The Academy separated tool use from awardable authorship; a clinical study narrowed its claim to text records; harness and memory papers moved performance into system architecture; creative judgments varied by production phase; and Anthropic's reported valuation remained a financing negotiation, not a completed market mark.
1. The Academy makes human authorship a condition for acting and writing awards
The Academy of Motion Picture Arts and Sciences said eligible performances must be credited, performed by humans, and made with their consent. Screenplays must be human-authored, and the organization reserved the right to request details about a production's AI use and authorship.
The rule permits generative tools while reserving acting and writing recognition for people. Awards eligibility now turns production records into evidence: consent, credited performers, script provenance, and AI-assisted revisions must preserve a human-authorship chain that the Academy can inspect.
Sources: AI-generated actors and scripts are now ineligible for Oscars · Awards rules approved for the 99th Oscars
2. Anthropic's reported private round targets a $900 billion valuation
Anthropic was seeking roughly $50 billion at a valuation around $900 billion, with investor allocations due quickly and a possible close within two weeks, according to TechCrunch. The report also cited company revenue-run-rate claims and unnamed sources familiar with the financing.
Anthropic declined to comment, so the amount, timing, valuation, and revenue figures were not confirmed transactions. The report is best read as evidence of extraordinary investor demand and compute financing needs, not as a completed mark for the company or the wider AI market.
Sources: Anthropic potential $900 billion valuation round could happen within two weeks · Anthropic weighs a new funding round
3. OpenAI's o1 performs strongly on a small emergency-care study
Researchers tested o1 and GPT-4o against two internal-medicine attending physicians using records from 76 emergency-department patients. At initial triage, o1 produced an exact or close diagnosis in 67% of cases, compared with 55% and 50% for the two physicians.
The comparators were internal-medicine attendings rather than emergency physicians, and the text records omitted visual and interpersonal information available in care. No patient outcome was measured. The 67% result supports a prospective decision-support trial with clinicians in the loop, not autonomous diagnosis from an emergency chart.
Sources: Harvard study compares AI diagnoses with two physicians · AI model performed strongly on real patient records
4. A preprint treats coding-agent harnesses as an optimization target
Agentic Harness Engineering, first submitted April 28, describes a closed loop that represents editable harness components explicitly, compresses execution trajectories into usable evidence, and checks each proposed edit against later task outcomes. The method targets tools, middleware, memory, and control logic rather than only prompts or models.
The accessible paper was revised after this issue date, so later headline scores cannot be treated as contemporaneous May 3 evidence. The durable engineering proposition is that observability and reversible harness changes can be evaluated independently from a model upgrade.
Sources: Agentic Harness Engineering preprint, version 1
5. Schema-grounded memory outperforms retrieval baselines in a new preprint
A May 1 revision of the xmemory preprint argues that persistent AI agent memory should behave like a system of record. Its write path detects objects and fields, extracts values, validates them, and stores structured records so reads do not repeatedly infer facts from retrieved prose.
The authors report 97.10% F1 on their end-to-end memory benchmark versus 80.16% to 87.24% for third-party baselines. These are author-run evaluations of a proposed system, but they support testing updates, deletions, unknown values, and relational queries separately from semantic recall.
Sources: Schema-grounded memory preprint, version 2
6. Creative benchmark finds no model wins every production phase
Contra Labs collected 5,940 pairwise judgments, 5,940 scalar ratings, and 3,675 written responses across 93 prompts and five creative domains. Evaluators assessed ideation, mockup, and refinement, and no tested model led all three phases within any domain.
Contra selected the evaluator group and reviewed the prompts, so the rankings describe its study rather than a universal creative hierarchy. Treating disagreement as signal is the sharper contribution: adherence and usability can converge while taste remains legitimately plural, making phase-level score distributions more informative than one winner.
Sources: The Human Creativity Benchmark · CrowdTruth 2.0 methodological foundation, version 1