Three evaluations expose gaps between nominal capability and usable performance. Dense inputs shrink effective context, Claude's chemistry gains stay bounded by a small test set, and Meta's withdrawn personalized feed shows how relevance can amplify fabricated material when provenance disappears.
1. Lexical density shrinks the effective context window
Researchers tested open-weight LLMs from 9 billion to 685 billion parameters on three retrieval tasks held near 12,000 tokens. Models that were almost perfect on sparse inputs fell below a 60% retrieval score when the same-length context carried more distinct information.
The preprint isolates lexical density from length and answer position, then finds recovery when density falls. The same 12,000-token budget therefore handles repetitive transcripts better than contracts, logs, or technical records packed with unique facts.
Sources: Dense Contexts Are Hard paper
2. Claude's NMR study shows promise within a narrow chemistry test
Anthropic compared three Claude models with ChemDraw and MestReNova on 20 post-cutoff compounds. It reports that Opus 4.7 had a mean hydrogen-shift error of 0.079 ppm and recovered all eight simpler inverse structures on every attempt from formula and one-dimensional NMR spectra.
The evaluation covers four scaffold families, three solvents, and 15 inverse problems, with starting-material hints on harder cases. Its strongest scores locate a narrow spectral-assistance capability while leaving broader blind chemical identification unresolved.
Sources: Anthropic's Claude chemistry study
3. Meta withdraws an AI-generated personalized content feed
Meta tested a For You section in its standalone AI app that generated article-like text and images from personalized prompts. The Verge found absent sourcing, apparent fabrications, distorted public figures, and no obvious AI label; Meta said the limited test would be deprecated after questions from the publication.
Personalization made unsourced output appear tailored and therefore more credible to each reader. Removing the feed after outside scrutiny exposes a product-review gap around citations, labeling, regeneration, and depictions of real people.
Sources: The Verge's examination of Meta's generated feed · Meta's AI app announcement