Supply Chain AI Index

A benchmark for AI models on real supply chain work.

The Supply Chain AI Index will evaluate how well AI models answer supply chain questions, read supply chain documents and carry out agentic workflows, as established benchmarks do for coding and legal work. The first edition has not been published.

First editionComing soon

What it will evaluate

  • Supply chain questionsAnswers to the questions teams ask every day
  • Documents and dataReading the documents work depends on
  • Agentic workflowsMulti-step tasks carried out with tools
  • Grounding and safetyCiting sources and admitting gaps

No results yet. Scores will appear here only when the evaluation is complete.

About the Index

A benchmark for supply chain work.

General benchmarks show how models write, reason or code. None shows how a model handles a delayed shipment, a disputed invoice or a supplier that confirmed only part of an order. The Index is planned to measure exactly that.

On this page

Everything we can state today about the Index.

  • What the Index will evaluate, in four areas
  • How we plan to build the tasks, score the models and publish the results
  • How supply chain teams and model providers can contribute

Not here yet

We will add each of these when it exists.

  • Scores, rankings or results of any kind
  • The task set and the list of models evaluated
  • A publication date for the first edition

If your team runs planning, procurement, logistics, warehousing or finance operations, the questions and workflows you handle every day can become tasks in the benchmark. If you build models, you can ask for yours to be included.

What the Index evaluates

Four areas, each scored on its own.

A single score hides very different strengths. A model can answer domain questions well and still fail a multi-step workflow, so the Index will report each area separately.

  • Area 01

    Supply chain questions

    How well a model answers planning, procurement, logistics and finance questions, scored against reference answers written by practitioners.

    PlanningProcurementLogistics
  • Area 02

    Documents and data

    How accurately a model reads purchase orders, invoices, proofs of delivery and contracts, and extracts the fields a workflow needs.

    InvoicesProofs of deliveryContracts
  • Area 03

    Agentic workflows

    How reliably a model completes multi-step tasks with tools: finding the records, comparing them, proposing an action and stopping where a person decides.

    Tool useMulti-stepApprovals
  • Area 04

    Grounding and safety

    Whether a model cites the records behind its answer, says when data is missing instead of guessing, and stays within the permissions it is given.

    CitationsMissing dataPermissions

Contribute

Help shape the first edition.

Contributing starts with one structured conversation. You share the kind of work your team does, and we turn it into anonymised tasks.

  • Supply chain teams can contribute anonymised questions, documents and workflows
  • Model providers can ask for their models to be included
  • Contributions are anonymised by default. Your organisation is named only with written permission
  • No cost, no purchase requirement and no obligation to be a customer
  • We plan to share the results with contributors before they are published

How contributing works

  1. Write to usUse the contact page and mention the Index. Tell us whether you run supply chain operations or build models.
  2. One structured sessionWe agree what you contribute, such as typical questions, documents or workflows, and how it is anonymised.
  3. Your work becomes tasksContributions are reviewed, anonymised and used only to evaluate models in the benchmark.
Contribute

Nothing is collected on this site. Contributing starts with a conversation.

Questions

What contributors ask before they start.

If your question is not here, write to us through the contact page and mention the Index.

No. The first edition has not been published, and this page contains no results. We will publish scores only when the evaluation is complete, together with the task categories, the scoring method and the model versions tested.

How well AI models answer supply chain questions, read supply chain documents and carry out multi-step agentic workflows, and whether they cite their sources and stop where a person should decide. It measures models, not companies.

If they are included, they will run the same tasks under the same rules as every other model, and the results will say which models are ours.

No. Any team that runs supply chain operations can contribute, and any model provider can ask for its models to be included. There is no cost or purchase requirement.

No. Contributions are anonymised before they become tasks, and the original documents are never published or resold. We plan to publish the task categories and examples, not your data.

Help build the benchmark.

Share the questions and workflows your team handles, or ask for your model to be included. There is no cost and no obligation to be a customer.