How we collect.
Last updated 19 August 2026
What factget retrieves, what it refuses to retrieve, and what happens to your data once it's in a pipeline. Stated plainly, so you can bring this to whoever needs to sign off before we talk.
factget retrieves from public, logged-out sources only. Retrieval is read-only: we don’t submit forms, post content, complete transactions, or modify anything at the source.
robots.txt is respected. Machine-readable rights reservations are honoured. Retrieval is rate-limited rather than run at whatever speed a source happens to tolerate.
There’s a class of source we won’t touch, regardless of how the request is framed:
- Sources behind a login, a paywall, or a terms-of-service gate.
- Anything that requires credentials, a shared account, or an account we create ourselves.
- Circumventing a site's technical access controls.
- Personal data as a primary collection target.
- One-off, bespoke scrapes outside a defined, recurring schema.
Personal data is avoided by default, enforced at the schema level rather than left to a reviewer to catch after the fact.
Where a dataset unavoidably includes personal data, your organisation is the data controller and factget acts as the processor. A data processing agreement is available on request.
| Mode | What it means |
|---|---|
| Managed | factget hosts the pipeline and the dataset. |
| No-retention (ZDR) | Extraction runs without factget retaining a copy once the dataset is delivered. |
| Customer VPC | The pipeline runs inside your own cloud environment. |
In every mode, your material is not used as training data, by factget or by any subprocessor. Export is continuous rather than a one-time handoff, so leaving with everything you have is always available, not a negotiation.
| Provider | Purpose | What it touches |
|---|---|---|
| Model provider | Source investigation and extractor generation | The source content being investigated |
| Proxy / unblocking provider | Routing retrieval requests | Request metadata, not the extracted data itself |
| Cloud host | Infrastructure and storage | Pipeline infrastructure and, in managed mode, the dataset |
This list changes as the product does. Customers on a contract are notified before a new subprocessor is added.
Every value in a dataset carries three things: the source it came from, the specific path that reached it (page, selector, or field), and when it was retrieved. That trail is built while the extraction runs, not reconstructed afterwards.
It’s a record of where a value came from, not a certification against any external standard.
factget does not hold a SOC 2 report today. A Type I engagement begins once a customer’s timeline requires it.
In the meantime, the controls available are architectural rather than a badge: deployment in your VPC or in no-retention mode, a collection policy limited to public, logged-out sources, and continuous data portability.
Running a security review? Send the questionnaire to apurv@factget.ai and we’ll answer it directly.