Transparency

What we publish and why

Transparency is not a legal obligation we fulfil minimally. It is a design principle. We publish what we know, including what we do not know and what we got wrong.

Publications

What we make public

Model Cards

Every model we ship has a published model card documenting its intended use, evaluation results, known limitations, and out-of-scope uses. Model cards are updated with each significant checkpoint. They are written to be readable by non-technical stakeholders, not just researchers.

Training Data Provenance

We publish a high-level data provenance report for each major model, covering the categories of data used, the sources and their licensing status, and the filtering and curation methodology. We do not publish raw datasets, but we publish enough to allow meaningful scrutiny of our data practices.

Evaluation Results

We publish benchmark results on standard academic evaluations and our own internal evaluation suite, including low-resource and multilingual benchmarks that most labs do not report. We include disaggregated results across language, domain, and demographic group where available. We do not cherry-pick benchmarks.

Known Limitations

We maintain and publish a known limitations register for each deployed model, documenting failure modes we have identified, the conditions under which they occur, and what mitigations we have or have not yet deployed. This register is updated on a rolling basis as new limitations are discovered.

Reasoning

Decisions we've made and why

We believe that explaining the reasoning behind significant decisions is as important as the decisions themselves. Where we have made choices that have meaningful implications for users, researchers, or society, we explain why we made them and what alternatives we considered.

This includes decisions to restrict certain capabilities, to refuse certain deployment contexts, to prioritise certain languages in evaluation, and to set our Responsible Scaling Policy thresholds where we did. These explanations are published in the transparency log below and updated as decisions evolve.

Log

Transparency log

A dated record of significant decisions, policy changes, and disclosures.

April 2026

Policy

Sattam.ai access restricted to licensed legal practitioners

We decided to limit early access to Sattam.ai to licensed advocates and law clerks rather than making it broadly available. Our reasoning: legal information provided to unlicensed users without appropriate context creates a meaningful risk of harm in high-stakes legal situations. We will expand access as we develop appropriate safeguards for non-practitioner use.

February 2026

Safety

Clark CL-2 safety evaluation results published

We published full results from our CL-2 capability and safety evaluations, including red-team findings and the two areas where Clark performed below our internal threshold on first assessment: persuasion resistance in regional language contexts and factual accuracy on contested historical events. We describe the mitigations applied and their measured effectiveness.

January 2026

Governance

AI Constitution v1.0 ratified, public comment period summary published

After a 90-day public comment period, we ratified Version 1.0 of our AI Constitution. We received 47 substantive comments from researchers, civil society groups, and legal scholars. We publish a full summary of comments received, our response to each, and where comments led us to revise the document. Three comments led to material changes.

November 2025

Data

Training data audit: three sources removed after licensing review

An audit of Clark's pre-training corpus identified three data sources whose licensing terms we concluded did not clearly permit use in commercial AI training. We removed those sources from the corpus and conducted a targeted fine-tuning pass to reduce any residual influence. We publish the names of the removed sources and our legal reasoning.