Skip to main content

Title and land records

Title plant infrastructure, built for county records.

Extraction, plant construction and court record ingestion for title plant operators, search firms and underwriters. Purpose-built for how counties actually record, not adapted from generic document AI.

100,000+
documents per day
All US counties
all US states
40–60%
less manual data entry

The work that doesn't scale

Title plants hold decades of deeds, mortgages, liens, assignments and court records across more than 3,000 counties, and most of it arrives as scanned images. Turning those images into searchable, indexed data has historically meant people reading documents and typing.

That model is under pressure from three directions at once. The experienced courthouse abstractor workforce is retiring without replacement. County digitisation costs more than most independent operators can carry. And lenders increasingly want title intelligence through an API rather than a research report.

Generic document AI doesn't close the gap, because courthouse records aren't generic documents.

Document extraction and indexing

What Leo reads

Leo ingests multi-page TIF batches — the standard county export format — auto-detects and corrects rotated pages, and applies OCR, NLP and machine learning to extract structured fields. Output lands in SQL Server and surfaces in a review interface where trained reviewers verify and approve before export.

Instrument
WARRANTY DEED · Filed 08/14/2026 · Instrument #2619330
  • Type: Warranty Deed
  • File date: 2026-08-14
  • Instrument: 2619330
Parties
… a married woman as her sole and separate property, and … a married man as his sole and separate property, grant(s) to … a single man …
  • Grantor ×2 · married
  • Grantee ×1 · single
  • Surname-first normalised
Legal description
SE1/4SW1/4 of Section 14, T.22S., R.2E., of the N.M.P.M. of the U.S.G.L.O. Surveys, Dona Ana County, New Mexico
  • Quarter: SE1/4 SW1/4
  • Section: 14
  • Township: 22S
  • Range: 2E
  • Meridian: N.M.P.M.
  • Survey: U.S.G.L.O.
Subdivision
Lot 19 in Block M of Del Prado Subdivision Phase 5, as shown on the plat filed in Plat Book 22, Folio 584-586
  • Lot: 19
  • Block: M
  • Subdivision: Del Prado
  • Phase: 5
  • Plat book: 22
  • Folio: 584-586
Oil and gas lease
… and 14 other lessors, whose interests are set out in Exhibit A attached hereto, to … as Lessee, covering the lands described …
  • Lessor ×15
  • Lessee ×1
  • Exhibit A parsed
  • Multi-party resolved

Raw recorded text above, structured fields below. Every field traces back to its source page.

Instrument types

Warranty, quitclaim, correction and trustee deeds. Oil and gas leases and lease amendments. Assignments and partial assignments. Lis Pendens filings with prose-style party lists. Affidavits of heirship, plat dedications, mineral conveyances. Civil case records with nine party role types.

The review interface

Leo's review screen for a warranty deed: extracted identification, date and party fields on the left, the recorded instrument on the right.
Extracted fields sit beside the recorded instrument. Page badges show which page each field came from, so a reviewer can check any value against its source without leaving the screen. Party details redacted.
Leo's review screen for a thirteen-page deed of trust, showing the page strip used to move between pages of a single instrument.
Multi-page instruments carry a page strip. A thirteen-page deed of trust is reviewed as one document, with fields drawn from whichever page holds them.

Building a county plant from raw data

Counties deliver instrument records as CSV exports — unvalidated, inconsistently formatted, and structured differently in every jurisdiction. Turning that into a production title plant database is a data engineering problem, not a data entry problem.

A three-phase pipeline

Prep ingests and classifies raw county data, parsing every recorded instrument and categorising legal descriptions by type. Build assembles cleaned data into a relational title plant schema, resolving parties, volumes, pages, record types and document references. Wrap applies county-specific configuration, corrects legal strings, registers to the production server and runs validation.

Human review gates, not blind automation

Two structured review gates are built into the pipeline: one for records that can't be auto-classified, one for party name strings that exceed column limits. A person decides; the system records the decision. Automated milestone backups are taken at each phase boundary.

Every county build improves the next

Abbreviation rules and county-specific overrides live in a central table with audit timestamps. Knowledge accumulated building one county is available to the next, which is what makes "all US counties" a configuration problem rather than a rebuild.

Federal court records, ingested weekly

Active bankruptcy cases create automatic stays and other encumbrances affecting ownership and transfer rights. Most plants either monitor PACER by hand or buy a commercial bankruptcy feed.

Our pipeline connects to PACER on a configurable weekly schedule, retrieves new filings across configured jurisdictions for Chapters 7, 11, 12 and 13, pulls complete case detail including debtors, attorneys, trustees and creditors, normalises party names against a configurable rules engine, and inserts structured instruments directly into your plant.

Nothing commits without approval

Before any normalised party name reaches your production database, a designated operator reviews and approves the batch. That produces an auditable approval trail for every ingestion cycle — and means the data entering your plant has been seen by someone who understands it.

The service runs as a Windows background service handling PACER authentication, XML parsing, SQL import and notifications. Staff access a desktop application with a live pipeline dashboard and a normalisation rules manager, over VPN or Remote Desktop, with credentials protected by Windows DPAPI.

Built for county scale

Throughput

Batch processing with automated checkpoints and structured review gates, built into the pipeline rather than bolted on. 100,000+ documents per day, with 5–10× faster processing than manual review and entry.

All US counties, all US states

Every county records differently — different formats, different abbreviations, different conventions for party names and legal descriptions. Leo handles a new county as configuration, not a rebuild. County-specific rules are held centrally and improve every build that follows.

Human review where it matters

Confidence scoring flags exceptions automatically. Trained reviewers verify and approve before anything reaches a production plant, and every approval is recorded with full lineage — source file, extraction history, reviewer, export.

Who we work with

Title plant operators

Building or refreshing county property record databases, and under pressure to digitise without the capital to do it by hand.

Title search and abstracting firms

Processing raw county feeds into searchable archives, and looking to move off labour-intensive margins toward AI-assisted ones.

Title insurance underwriters

Needing normalised, validated, auditable property data for risk assessment, with lineage that stands up to scrutiny.

Data providers and technology firms

Building title plant infrastructure for regional or national coverage.

Start with a county

Send us a sample batch from a county you're working. We'll extract it, show you the output, and you can judge the results against your own data.