Wire Observer.
Technology

Anthropic Pulls Public Model Evaluation Datasets, Citing Temporary Safeguard

Anthropic Pulls Public Model Evaluation Datasets, Citing Temporary Safeguard

Anthropic, the AI startup behind the Claude series of language models, announced on Tuesday that it is withdrawing its publicly hosted model evaluation datasets from the internet, describing the step as a short‑term measure rather than a permanent shutdown.

The evaluation sets, which include benchmark scores, prompts, and performance logs, have been accessible to researchers and developers for several months, providing a rare glimpse into the capabilities and safety characteristics of Anthropic's models. By making this data openly available, the company previously aimed to foster transparency and enable third‑party verification of its claims.

In a brief statement, Anthropic said the removal is driven by emerging concerns over potential misuse of the evaluation material and the competitive landscape of large‑language‑model development. The firm indicated that keeping the data online could inadvertently aid actors seeking to reverse‑engineer or exploit model behaviors, especially as the field accelerates toward more powerful systems.

The decision has sparked discussion within the AI research community about the balance between openness and security. Scholars rely on shared benchmarks to compare progress, reproduce results, and identify safety gaps. With Anthropic’s data temporarily offline, some researchers warn that the move may slow collaborative efforts and limit independent assessment of the company's safety claims.

Anthropic did not specify a timeline for reinstating the datasets, but it emphasized that the withdrawal is not permanent. The company indicated it is exploring alternative ways to share evaluation results that mitigate risks while preserving scientific rigor. Observers note that similar restrictions have appeared at other leading AI firms, suggesting a broader trend toward more guarded dissemination of model performance information as the industry grapples with both innovation and responsibility.

Source: Gizmodo
Diya Sharma — AI & research desk.

Comments (0)

Be the first to comment.

Join the discussion

Protected by reCAPTCHA v3

Related