# A model whose full recipe is visible.

Ai2 released OLMo 2 32B, a language model accompanied by publicly available data, code, weights, training details, and a reproducible training recipe.

Canonical: https://brightaifuture.com/discoveries/olmo-2-32b
Format: discovery
Source publication: 2025-03-13
Bright publication: 2026-09-07
Substantive update: None recorded
Evidence and review: Demonstrated; confidence: unassessed; approved; ai-assisted. AI-assisted editorial comparison with the cited primary source; result, setting, source date and limitations retained. Independently checked within the research team. Publication authorized by the site owner; no human source review is claimed.

## The human problem

Researchers and smaller organizations can find it difficult to inspect, reproduce, or adapt capable language-model systems when only a hosted interface or weights are available.

## The prior constraint

Many model releases did not make the end-to-end development pipeline available for scrutiny and reuse.

## AI’s actual role

A 32-billion-parameter language model trained to 6 trillion tokens and post-trained with Tulu 3.1.

## The documented result

Ai2 reports that OLMo 2 32B outperformed GPT-3.5 Turbo and GPT-4o mini on its selected multi-skill academic benchmark suite. Ai2 also states that all ingredients of its end-to-end training recipe are available.

## Why it may matter

People can inspect and adapt more of the system behind a model, which may make research, auditing, and specialized local development more practical.

## Limitations

The performance comparison is the developer's benchmark result, not evidence of better outcomes in workplaces or communities.

Availability of a recipe does not remove the substantial compute and expertise required to train or fine-tune a model.

This source does not establish performance, safety, or fairness across all languages and uses.

## Unresolved questions

Can independent groups reproduce the reported results?

How does the model perform and fail in domain-specific and non-English work?

What safety and bias findings emerge when others adapt the recipe?

## Provenance and history

{
  "dates": {
    "eventDate": null,
    "publicationDate": "2025-03-13",
    "captureDate": "2026-09-07",
    "lastReviewedDate": "2026-09-07"
  },
  "provenance": {
    "origin": "editorial",
    "externalId": "https://allenai.org/blog/olmo2-32b"
  },
  "revisions": [
    {
      "id": "revision:9e316fcad06b5346be4c",
      "recordedAt": "2026-09-07",
      "summary": "People can inspect and adapt more of the system behind a model, which may make research, auditing, and specialized local development more practical.",
      "sourceIds": [
        "source-olmo-2-32b"
      ]
    }
  ],
  "corrections": []
}

## Original sources

- [OLMo 2 32B: First fully open model to outperform GPT 3.5 and GPT 4o mini](https://allenai.org/blog/olmo2-32b)

## Continue exploring

- [Open intelligence](https://brightaifuture.com/worlds/open)
- [What changes when powerful models become open-weight?](https://brightaifuture.com/threads/open)
- [Someone builds on it](https://brightaifuture.com/open-intelligence)
