Overview
Disclosure is the common thread today. Cursor acknowledged the open model beneath Composer 2, a game developer began removing AI assets that were never meant to ship, and new research argues that subjective annotations should preserve demographic disagreement instead of collapsing it into one label.
Developments
1. Cursor acknowledges Kimi beneath Composer 2
Cursor launched Composer 2 without initially identifying its base model, then acknowledged that it started from Moonshot AI's open Kimi K2.5. A Cursor executive said roughly one quarter of the final model's compute came from the base and the rest from Cursor's continued training; a cofounder called the missing attribution a mistake.
A derived model retains the base model's license and supply-chain dependencies even when additional training changes its behavior. Cursor's disclosure therefore belongs in model documentation alongside the training changes, hosting path, and evaluation scope that enterprise buyers use to assess legal and security exposure.
Sources: TechCrunch on Cursor's Kimi disclosure · Moonshot's Kimi K2.5 model card
2. Crimson Desert developer audits AI-generated assets
Pearl Abyss acknowledged that AI-generated art appeared in the released version of Crimson Desert, apologized for failing to disclose its use, and said the material had been intended as temporary. The studio announced a comprehensive audit to identify and replace remaining AI assets.
The incident is a content-supply-chain failure. Asset-level provenance and release checks can distinguish approved generated material from placeholders before thousands of files enter a production build; an intention to replace temporary art provides no control at release time.
Sources: The Verge on Crimson Desert's AI-art audit · Engadget on the developer's replacement plan
3. Subjective annotation research preserves demographic disagreement
A March 22 preprint introduces Perspective-Driven Inference for LLM-assisted annotation when politeness, offensiveness, or another subjective judgment varies across demographic groups. The method estimates a distribution of group responses and directs a limited human-labeling budget toward groups for which the model is least accurate.
Tests on politeness and offensiveness ratings improved estimates for harder-to-model groups relative to uniform human sampling, within the paper's selected datasets and group definitions. Its larger contribution is statistical: disagreement can be the quantity under study, so replacing it with one synthetic label can erase the population difference an analysis is meant to measure.
Sources: Multi-Perspective LLM Annotations preprint, version 1