Last run: July 2026 · Reproducible from our eval harness
Most "AI document extraction" accuracy claims are unverifiable. Ours isn't: we test against rent rolls filed with the U.S. Securities and Exchange Commission — public documents anyone can check — and publish exactly what we measured.
| Document | Type | Source filing | Result |
|---|---|---|---|
| Hartman portfolio rent roll (2018) | Multi-property tenant-level rent roll, 8-K loan exhibit | SEC 0001446687-18-000051 | PASS |
| Uptown Tower rent roll (2013) | Office unit/tenant rent roll, 8-K loan schedule | SEC 0000928953-16-000005 | PASS |
| Woolworth Building tenants | Tenant-level table from underwritten rent roll, CMBS prospectus | SEC 0000950136-05-003719 | PASS |
| Mazza GrandMarc unit mix | Student housing (232 units) unit-mix summary, CMBS prospectus | SEC 0001539497-13-001242 | PASS |
| Luxe at 1548 unit mix | Multifamily (54 units) unit-mix summary, CMBS prospectus | SEC 0001539497-15-001777 | PASS |
| Granada Gardens unit mix | Multifamily (940 units) unit-mix summary, CMBS prospectus | SEC 0001539497-14-001559 | PASS |
This benchmark shows exact-match performance on real, messy, public rent-roll documents — including OCR-degraded source text (one document literally spells "LIFESTYLES" as "LJFESTYLES"; we extract what's printed, faithfully). It does not claim every possible document will extract perfectly. That's why every extraction ships with a verification report: recomputed totals reconciled against stated totals, and any value the model can't read with confidence is flagged, never guessed. You always know which fields to double-check.
The corpus grows over time. Want your (anonymized) document format represented? Send it over.