Get the facts.
Keep them current
factget extracts structured data from websites, documents, PDFs, images and other difficult sources, then monitors the source and keeps the data up to date.
Build
turn difficult sources into a structured dataset
Extract
investigate once, then run deterministically at scale
Maintain
detect changes, validate results, repair extractors
From difficult sources
to usable datasets.
Define the information you need. factget finds it across your sources, structures it into a consistent schema, and keeps it current as those sources change.
One page is easy.
A million pages is infrastructure.
A dataset is rarely one page. It is thousands of pages, hundreds of sources, several formats, and a source that may change tomorrow.
Screen every name on our counterparty list against global sanctions registers.
- –flag only close matches, not every near-spelling
- –include the source list and match reason
- –keep results current as lists are updated
| entity | type | country | program | source |
|---|---|---|---|---|
| Entity | Russia | OFAC SDN | SDN List | |
| Individual | Turkey | EU Consolidated | EU OJ | |
| Entity | China | BIS Entity List | Fed Register | |
| Entity | UAE | UK OFSI | UK Gazette |
- entity
- ↳ OFAC SDN List · entry 28491
- program
- ↳ OFAC SDN List · program UKRAINE-EO13662
- country
- ↳ OFAC SDN List · address field
We detect two kinds
of change.
Getting the data is half the problem. Keeping it correct is the other half, and a dataset can go wrong in two entirely different ways.
The data changed.
update the dataset
- Source
- Q3 operations report
- Extractor
- unchanged · still valid
Production changed: 2.4M → 2.7M bbl/d
dataset updated · no revisit required
The source changed.
repair the extractor
.product > .price- Value
- may be identical
- Extractor
- invalid
- Running
- Source changed
- Anomaly detected
- Reinvestigate
- Repair extractor
- Validate
- Resume
Fig. 02A moved value and a moved selector are different failures
Websites change.
Your dataset shouldn’t break.
Don’t use an LLM100,000 times.
LLMs are excellent at understanding unfamiliar sources. They don’t need to rediscover the same extraction strategy for every page.
⋮
deterministic
⋮
Intelligence once.
Deterministic at scale.
- 05
Validate
checked against the schema it claims
- 06
Monitor
the source is re-read on a schedule
- 07
Change?
value moved, or structure moved
- 08
Repair → continue
rebuild the extractor, then resume
Same first cost: paid once, not per record.
- Lower marginal cost
- The expensive reasoning happens once instead of on every extraction.
- High throughput
- Deterministic execution avoids an inference cycle on every request.
- Predictable output
- The same extraction logic runs repeatedly across large datasets.
Don’t just extract.
Check it.
At one page you can eyeball the result. At a hundred thousand, the system has to be the thing that notices. Extraction that isn’t checked is just a plausible-looking guess at scale.
- Extract
- Validate
A failed check routes into the same reinvestigation path that handles a changed source.
- Schema
- fields, types and units are the ones the dataset promised
- Completeness
- a row that lost a value is flagged, not shipped half-empty
- Consistency
- values that contradict the rest of the column get surfaced
- Anomaly
- a figure that moves further than it should is held for review
- Provenance
- every value keeps a pointer back to where it was found
- No invented scores
- We publish accuracy figures when we have measured them, not before.
Any source.
One structured output.
The format is a detail. Every source resolves to the same contract, with a pointer back to where each value was found.
- Web
- Image
- Scan
- Table
- Spreadsheet
- API
- Language
- 01Retrieve
- 02OCR
- 03Understand
- 04Extract
- 05Translate
- 06Structure
- 07Deliver
- Provenance
- every value points back to where it was found
- Validation
- output checked against the schema it claims
- Reproducible
- same source and logic, same result
- Privacy
- your material is not anyone else's training data
- Deployment
- runs where your data is permitted to live
Built for teams whose input is someone else’s output.
Filings, reports, and live data — structured on arrival.
Financial teams track thousands of entities across regulatory filings, corporate disclosures, and pricing pages. factget turns fragmented sources into a single monitored dataset.
Sanctions screening
Continuously extract and normalise sanctions lists across jurisdictions into one canonical schema.
One canonical schema across every registry you track
Earnings extraction
Pull revenue, net income, segment breakdowns and footnotes from structured and semi-structured filings.
Structured financial data from unstructured PDFs
Pricing & rate monitoring
Track posted rates, fee schedules, and product terms across hundreds of financial institution pages.
An alert when a posted rate moves, not a quarterly re-check
Fund fact sheets
Extract NAV, holdings, performance tables, and risk metrics from monthly fact-sheet PDFs.
Tabular data from documents that change layout quarterly
Build a datasetfrom your sources.
Tell us the sources and the fields you need. We’ll show you what factget can build, and keep current.