Research Asset

The Starred Menu Corpus

What Italy’s three-star restaurants actually served, year by year, from 1999 to today. Not a snapshot of what is online now — a dated record of what changed.

Menu data is easy to collect once and almost impossible to reconstruct backwards. A restaurant that reworked its tasting menu last spring has erased the previous one: it survives only in web archives, under URLs that moved, in editions nobody labelled with a date. Most of the work behind this corpus is not collecting. It is dating.

15

three-star venues

213

dated menu versions

13,205

dishes

184

recovered from archives

1999–2026

years covered

As of August 2026. Updated by hand after each significant run.

Method

How it is built

One crawler, two backends: the live site and the Internet Archive, running the same heuristics. A page found today and a page archived in 2014 are not the same kind of observation, and the corpus does not pretend otherwise — every edition is dated by its own capture, never by the day we happened to find it.

Differences between versions are read the way a cook would read them. A dish that was reworded is recorded as reworded, not as one dish leaving and another arriving. The corpus re-runs itself every month; when nothing has changed, nothing is re-extracted.

Provenance

The rules that make it citable

A scrape becomes a corpus when you can say where every line came from and why it is shaped the way it is. These four rules are what separate the two.

Extraction is verbatim

A menu that mixes Italian and English is a finding, not a defect to clean up. What the restaurant wrote is what the corpus holds.

Language belongs to the edition

The Italian and the English version of one menu are two forms of one thing, paired by the day they were captured — not two unrelated documents.

Every dish carries its source

Each line holds the permanent address of the snapshot it was read from. Any claim made from this corpus can be checked against the page that produced it.

Absence is data

A restaurant that publishes no menu is recorded as publishing no menu. Silence is part of the record, not a gap in it.

Questions it can take

What a dated record makes answerable

  • How the price of a tasting menu moved across a decade, one venue at a time.
  • When an ingredient or a technique first enters the repertoire — and whether it spread from there.
  • Which dishes survive ten years, and which return every season.
  • Whether the Italian and English editions of the same menu are saying the same thing.

Limits

What it cannot answer yet

Stated plainly, because a corpus you cannot argue with is a corpus nobody can use.

Archive coverage is uneven

Some venues are documented from 1999; others begin in the 2010s. Certain years exist in one language only — that is what the archive contains, and the gaps are visible in the data rather than smoothed over.

The error rate is not yet measured

Extraction has not been scored against a gold sample. That measurement is planned. Until it exists, this is a research instrument, not a published dataset.

Three stars only, for now

The record covers Italian three-star venues. Extension to one- and two-star restaurants is in progress and will change every number on this page.

Access

Read access on request

The corpus is not published as a download. It is a research instrument and it travels with its provenance — so access goes to researchers and institutions who tell us what they want to ask of it. Say what you are working on and we will open a read-only view.

Or write to hello@foodtechbootcamp.com.