OpenAI links AI reasoning extraction campaign to people associated with Moonshot

The ChatGPT developer says thousands of accounts were involved in attempts to extract protected model reasoning during July, with a core cluster linked to people associated with Ki

By The Register

OpenAI says it disrupted a coordinated campaign designed to extract protected reasoning from its artificial intelligence models, with part of the activity attributed to people associated with Chinese AI company Moonshot AI.

The company said the earliest activity it identified began on 1 July, initially at low volumes before increasing sharply later in the month.

On 24 and 25 July, OpenAI recorded around 16,000 requests using what it described as a relevant extraction pattern from more than 4,000 users.

Its subsequent investigation identified related prompt activity across a wider cluster of more than 15,000 users, with the company saying the campaign had been fully disrupted by 28 July.

OpenAI stressed that those figures represent attempted extractions and do not establish that all of the attempts successfully recovered protected reasoning.

The company described the activity as consistent with “adversarial distillation”.

Model distillation is a broader technique in which the outputs of one AI system are used to help train or improve another model. OpenAI uses the term adversarial distillation for systematic, unauthorised attempts to reproduce or improve another model using protected outputs or reasoning.

According to OpenAI, the July campaign targeted what it calls protected reasoning — internal information generated while a model works through a problem that is not normally included in its final response to the user.

The company said those involved did not break its encryption, compromise a database or obtain direct access to stored user conversations.

Instead, it says model interactions were manipulated so that protected reasoning could be reproduced in forms visible to the requester.

OpenAI attributed what it called a core cluster of the activity to individuals associated with Moonshot AI, the Beijing-based developer of the Kimi family of AI models.

However, it said it remained unclear whether every operator detected during the period came from a single organisation or actor.

The company responded by banning accounts associated with the activity and strengthening controls around account creation, infrastructure and monitoring.

It also said information about the technique had been shared with other AI developers through the Frontier Model Forum and with relevant government partners.

OpenAI argues that large-scale extraction of model reasoning presents both commercial and security concerns because it could allow another developer to reproduce capabilities without investing in the same training process or adopting equivalent safeguards.

The allegation comes amid growing competition between US and Chinese AI developers and increased scrutiny of the ways in which companies use the outputs of rival systems to improve their own models.

Earlier in September, Anthropic separately accused Moonshot AI and other Chinese AI developers of carrying out large-scale distillation activity targeting its Claude models.

The broader debate is complicated by continuing disputes over how generative AI companies themselves obtained training material.

Developers including OpenAI have faced legal challenges and criticism over the use of copyrighted material in model training, while AI companies increasingly argue that large-scale extraction of their own proprietary model behaviour should be restricted.

Those are separate legal and technical questions, however, and OpenAI’s latest disclosure concerns alleged unauthorised attempts to extract protected model reasoning through its services rather than a conventional intrusion into company databases.

Moonshot AI had not issued a public response to OpenAI’s latest attribution in the sources reviewed at the time of publication.

Open article on Cheshire Today