Complex documents. Structured data. In seconds.

We extract structured data from complex documents, ready for your systems, workflows and AI agents.

Emails, scans and PDFs supported · API-first integration

From Facility Agent

Subject Executed facilities agreement

Please find the executed agreement attached. Commitments are as set out in Schedule 1.

📎 SFA_executed.pdf · 312 pages

From the attachment

Pages
312
Lenders
5
Total commitment
€250,000,000✓
Evidence
Schedule 1, p.214

Senior Facilities Agreement · p.214 of 312

Schedule 1 — The Original Lenders

Bank A60,000,000
Bank B50,000,000
Bank C50,000,000
Fund D45,000,000
Fund E45,000,000
Total250,000,000

Commitment schedule

Lenders
5
Currency
EUR
Total commitment
€250,000,000✓
Read from
Text layer, p.214

SCHEDULE 1 — COMMITMENTS

Lender A40,000,000
Lender B40,000,000
Lender C40,000,000
Total120,000,000

Commitment schedule

Page
Turned upright, straightened
Lenders
3
Total commitment
$120,000,000✓
Read from
Scanned page 88

Credit Agreement · p.152 of 359

Lender Commitments

Lender A65,333,333.33
Lender B65,333,333.33
Lender C65,333,333.33
Others84,000,000.00
Total280,000,000.00

Commitment schedule

Stated total
$280,000,000.00
Sum of lenders
$279,999,999.99
Warning
Totals don’t match!

Already used by

Where Herein fits

Herein extracts information from complex documents, validates the results and delivers the data in a format that works with your existing systems.

What comes in

  • Emails and their attachments
  • Long agreements, read in full
  • PDFs, digital or scanned
  • Financial agreements and their schedules

Herein

Scans, parses and checks

Where it goes

  • Your CRM and other systems
  • AI pipelines and agents
  • Reviewers, with the evidence for each value

We build bespoke integrations around your documents, data requirements and existing systems.

Talk to us about bespoke integrations

Read, parse, ledger

What Herein extracts depends on the document and on what you need from it: financial figures, clauses, party names. The process is the same every time.

  1. 01 · Read

    Every page becomes text, with the position of each character.

    Digital documents are read from the file’s own text. Scanned pages are first turned the right way up by a visual model, then converted to text.

    Pages too damaged to read with certainty are flagged for a person to review. If you would rather have a best-effort read, we can do that instead.

  2. 02 · Parse

    The text becomes the structure your systems need.

    Code rules extract what they cover; how much they cover depends on the document.

  3. 03 · Ledger

    A record of what was found, and what wasn’t.

    Each document comes with a ledger: what was extracted and where, what was expected but missing (such as totals that don’t add up), and any exceptions from reading or parsing. It is there for a person to check against.

Input

  • PDF senior-facilities-agreement.pdf digital
  • PDF revolving-credit-facility.pdf digital
  • PDF amendment-and-restatement.pdf scanned
  • …
Herein

Output

{
  "document": { "page_count": 312 },
  "facilities": [
    {
      "id": "facility-1",
      "total_commitment": { "amount": "250000000", "currency": "EUR" },
      "commitments": [
        { "lender_name": "Bank A",
          "amount": { "amount": "60000000", "currency": "EUR" } },
        …
      ]
    }
  ],
  "extraction": {
    "schema_version": "1.1",
    "evidence": [
      { "path": "facilities[0]", "page": 214,
        "snippet": "Bank A | 60,000,000" }
    ],
    "warnings": [],
    "ledger": {
      "gaps": [ { "page": 6, "kind": "percentage", "text": "1%" } ],
      "complete": false
    }
  }
}

Every value comes with the page and the words it was read from. An empty field means not found, never zero.

The API
Covers credit agreements today, returning everything it finds in the agreement. APIs for other document types, and narrower ones that use fewer tokens, are in development.
Pricing
You pay for the tokens you use. Digital documents use fewer than scans, and badly damaged scans use the most. Integrations start with a call about what you need.
Testing
We backtest against historical documents that vary by type, language and whether they were scanned, and against the documents we process, and adjust the engine on the results.
In your stack
Herein can run as a step in an existing pipeline, or as a tool an AI agent calls. Sign in to read the Developer Docs for the full schema.

How Herein compares

Manual review General-purpose AI tools Herein
How it reads A person reads each document A model reads the text you give it Pages become positioned words; code extracts the values, models handle what code cannot
Output Notes or a spreadsheet, laid out differently by each reviewer Whatever shape the prompt asks for A fixed schema, in a format that works with your existing systems
Checking A second reviewer You check the answer against the document Each value carries its page and words; totals are checked against the document
When unsure Depends on the reviewer Can return a plausible value that is not in the document Leaves the field empty and lists the gap
Volume Limited by the size of the team One document or conversation at a time, unless you build tooling around it Many documents per request through the API
Into your systems Re-keyed by hand Copied across, or code you write and maintain Integrated with your systems by us, or called through the API

Anyone who handles complex financial agreements.

Banks and lenders

Commitments read out of signed facility agreements and into your own systems.

Law firms

A bundle of agreements turned into structured data you can read at a glance instead of page by page.

Loan administrators

Whole portfolios processed in bulk, on the same schema every time.

If you handle other kinds of complex document, tell us what they are.

Talk to us about bespoke integrations

We connect Herein to where your documents arrive and the systems that use the data. Tell us what you handle and where it needs to go.