Complex documents. Structured data. In seconds.
We extract structured data from complex documents, ready for your systems, workflows and AI agents.
Emails, scans and PDFs supported · API-first integration
Please find the executed agreement attached. Commitments are as set out in Schedule 1.
📎 SFA_executed.pdf · 312 pages
From the attachment
- Pages
- 312
- Lenders
- 5
- Total commitment
- €250,000,000✓
- Evidence
- Schedule 1, p.214
Senior Facilities Agreement · p.214 of 312
Schedule 1 — The Original Lenders
| Bank A | 60,000,000 |
| Bank B | 50,000,000 |
| Bank C | 50,000,000 |
| Fund D | 45,000,000 |
| Fund E | 45,000,000 |
| Total | 250,000,000 |
Commitment schedule
- Lenders
- 5
- Currency
- EUR
- Total commitment
- €250,000,000✓
- Read from
- Text layer, p.214
SCHEDULE 1 — COMMITMENTS
| Lender A | 40,000,000 |
| Lender B | 40,000,000 |
| Lender C | 40,000,000 |
| Total | 120,000,000 |
Commitment schedule
- Page
- Turned upright, straightened
- Lenders
- 3
- Total commitment
- $120,000,000✓
- Read from
- Scanned page 88
Credit Agreement · p.152 of 359
Lender Commitments
| Lender A | 65,333,333.33 |
| Lender B | 65,333,333.33 |
| Lender C | 65,333,333.33 |
| Others | 84,000,000.00 |
| Total | 280,000,000.00 |
Commitment schedule
- Stated total
- $280,000,000.00
- Sum of lenders
- $279,999,999.99
- Warning
- Totals don’t match!
What it does
Where Herein fits
Herein extracts information from complex documents, validates the results and delivers the data in a format that works with your existing systems.
What comes in
- Emails and their attachments
- Long agreements, read in full
- PDFs, digital or scanned
- Financial agreements and their schedules
Herein
Scans, parses and checks
Where it goes
- Your CRM and other systems
- AI pipelines and agents
- Reviewers, with the evidence for each value
We build bespoke integrations around your documents, data requirements and existing systems.
Talk to us about bespoke integrationsHow it works
Read, parse, ledger
What Herein extracts depends on the document and on what you need from it: financial figures, clauses, party names. The process is the same every time.
-
01 · Read
Every page becomes text, with the position of each character.
Digital PDFs are read directly from their text layer, preserving the position of every word. Scanned pages are first turned the right way up by a visual model, then converted to text.
Pages too damaged to read with certainty are flagged for a person to review. If you would rather have a best-effort read, we can do that instead.
-
02 · Parse
The text becomes the structure your systems need.
Code rules extract what they cover; how much they cover depends on the document. Anything left can be passed to models benchmarked against known answers. Model outputs are accepted only when their supporting evidence can be verified against the document.
-
03 · Ledger
A record of what was found, and what wasn’t.
Each document comes with a ledger: what was extracted and where, what was expected but missing (such as totals that don’t add up), and any exceptions from reading or parsing. It is there for a person to check against.
Input
- PDF senior-facilities-agreement.pdf digital
- PDF revolving-credit-facility.pdf digital
- PDF amendment-and-restatement.pdf scanned
- …
Output
{
"document": { "page_count": 312 },
"facilities": [
{
"id": "facility-1",
"total_commitment": { "amount": "250000000", "currency": "EUR" },
"commitments": [
{ "lender_name": "Bank A",
"amount": { "amount": "60000000", "currency": "EUR" } },
…
]
}
],
"extraction": {
"schema_version": "1.1",
"evidence": [
{ "path": "facilities[0]", "page": 214,
"snippet": "Bank A | 60,000,000" }
],
"warnings": [],
"ledger": {
"gaps": [ { "page": 6, "kind": "percentage", "text": "1%" } ],
"complete": false
}
}
}
- The API
- Covers credit agreements today, returning everything it finds in the agreement. APIs for other document types are in development.
- Pricing
- Pricing is based on document volume, workflow and integration requirements.
- Testing
- We backtest against historical documents that vary by type, language and whether they were scanned, and against the documents we process, and adjust the engine on the results.
- In your stack
- Herein can run as a step in an existing pipeline, or as a tool an AI agent calls. Sign in to read the Developer Docs for the full schema.
Compare
How Herein compares
| Manual review | General-purpose AI tools | Herein | |
|---|---|---|---|
| How it reads | A person reads each document | A model reads the text you give it | Pages become positioned words; code extracts the values, models handle what code cannot |
| Output | Notes or a spreadsheet, laid out differently by each reviewer | Whatever shape the prompt asks for | A fixed, versioned schema, in a format that works with your existing systems |
| Checking | A second reviewer | You check the answer against the document | Each value carries its page and words; totals are checked against the document |
| When unsure | Depends on the reviewer | Can return a plausible value that is not in the document | Leaves the field empty and lists the gap |
| Volume | Limited by the size of the team | One document or conversation at a time, unless you build tooling around it | Many documents per request through the API |
| Into your systems | Re-keyed by hand | Copied across, or code you write and maintain | Integrated with your systems by us, or called through the API |
Who it's for
Anyone who handles complex financial agreements.
Banks and lenders
Commitments read out of signed facility agreements and into your own systems.
Law firms
A bundle of agreements turned into structured data you can read at a glance instead of page by page.
Loan administrators
Whole portfolios processed in bulk, on the same schema every time.
If you handle other kinds of complex document, tell us what they are.
Talk to us about bespoke integrations
We connect Herein to where your documents arrive and the systems that use the data. Tell us what you handle and where it needs to go.