AI Agents

Multi-Agent Contract Review for Believe

A multi-agent LLM pipeline that reads full artist contracts in French, English and German and produces the same structured review a legal reviewer would: deal type, counterparts, royalty terms and breaches of 70 operating rules, with every output verified by a reviewer.

Multi-Agent Contract Review for Believe
  • 98.2% classification
  • 91-93% field concordance
  • FR · EN · DE

The problem

Believe is a music company that acquires labels and catalogues. Every artist contract in an acquisition has to be read before integration: what the deal is, who the counterparts are, what royalty rates apply, and whether the terms can be operated in Believe's systems at all.

That adds up to hundreds of contracts a year, and the reading was done by two people at about five contracts a day each. Contracts arrive in French, English and German, run long, and put the terms that matter in different places from one label to the next. Reading them is skilled work, and it was the bottleneck between signing an acquisition and running the catalogue.

Our approach

We built a multi-agent LLM pipeline that reads a full legal contract and produces the same output a legal reviewer would, in all three languages. The mission was delivered through Molia.

Specialised agents in a state graph. Instead of asking one prompt to do everything, the work is split across agents that run in a LangGraph state graph. One agent classifies the deal. Others extract the counterparts and the royalty terms into a typed schema. Another checks the contract against 70 operating rules and flags where it breaks them. Each agent has a narrow job, its own instructions and its own evaluation, and the graph carries the shared state between them.

Schema-constrained generation. Every extraction is generated against a Pydantic schema, so the output is typed and the model cannot silently return empty fields. A targeted verification pass then re-examines the cases most likely to be wrong, rather than paying to re-check everything.

Evaluation before modeling. Before any model work, we built a multi-metric benchmark with an error typology and measured inter-annotator agreement between human reviewers. That process exposed problems in the existing ground truth and in the function used to match model output against it, and both were corrected first. Without that step, every later accuracy number would have measured the benchmark's errors as much as the model's.

Reviewers stay in the loop. Every output remains reviewer-verified: the pipeline does the reading, people sign off. We shipped a standalone review application that four business reviewers now use in production. Their work feeds the loop that improves the system: review, correct, turn corrections into examples, re-benchmark.

This is the kind of document-processing agent our AI agents practice builds, with the model and evaluation work of our LLM engineering practice behind it.

Results

  • Contract type classification from 40% to 98.2%. Identifying what kind of deal a contract is decides which terms and rules apply, so it gates everything downstream.
  • 91-93% field-level concordance against reviewer ground truth for the extracted fields.
  • French, English and German contracts handled by the same pipeline.
  • In production with four business reviewers through a standalone review application.
  • A path to MVP. We delivered the architecture for the move from proof of concept to MVP, with technical documentation and a methodology note for the internal team taking it over.

Concordance is measured field by field against what a reviewer would have written, using the corrected ground truth and matching function. That is a stricter test than whether the output looks plausible, and it is the number the review team can reason about when deciding how much checking each field still needs.

What it takes to deploy

Contract review with LLMs is mostly an evaluation and workflow problem. The model calls are the easy part. Getting to a system a legal or business team trusts typically involves:

  • Fixing the ground truth first. Existing annotations are rarely consistent enough to benchmark against. Measuring agreement between reviewers shows how good "correct" can be before you ask a model to match it.
  • A schema the business owns. The fields, their types and what counts as "not stated in the contract" need to be agreed with the reviewers, because that schema becomes the contract between the model and downstream systems.
  • Operating rules written as checks. Rules such as the 70 used here have to be explicit enough for an agent to apply and for a reviewer to verify.
  • A review tool, not just a model. Reviewers need to see the extraction next to the source text, correct it quickly and have those corrections flow back into the evaluation set.
  • A handover plan. Documentation and a methodology note let the internal team extend the system and re-run the benchmark without the original builders.

We cover the evaluation side in depth in LLM contract extraction: build the evaluation first.

Where this applies

The same pattern fits any organization that has to read many long agreements before it can act on them: catalogue and rights acquisitions, licensing, supplier and procurement contracts, leases, and loan or credit agreements in banking & finance. Wherever the terms must be extracted into a system and checked against internal rules, a multi-agent pipeline with a reviewer in the loop can take over the reading.

Have a backlog of contracts that people are reading by hand? Talk to our team.

Capabilities & industries

More projects