BORME API

Product guide · Bulk data

Three million rows are easy to select. A reliable data product is harder.

2026-08-19 · 6 min · production boundary

By BORMEAPI Engineering

Exporting a table is easy in a development database. Exporting a multi-million- row company snapshot from a service that must keep answering normal API traffic is a different job.

BORMEAPI's ordinary database role has a short query ceiling. The full-company dump legitimately runs longer because it reads the maintained company view derived beforehand from processed BORME-A history, joins available NIF mappings, orders the result, encodes CSV and compresses the stream. Weakening every request to accommodate that one workload would turn a product feature into a reliability risk.

The Apify Actor exports selected publication events for a requested date range. The Hosted API Pro/Risk company dump exports the maintained, derived company snapshot. They are different datasets for different jobs.

The exception must be narrower than the guardrail

A safe bulk path gets its own disposable database session, known query and finite timeout. It does not borrow a normal low-latency pool connection, and its longer allowance disappears when the session closes. A disconnected client must also release the database work instead of leaving an orphan scan behind.

The exact timeout is not the product. The product is the complete boundary: which statement may run longer, who may call it, how often, on which connection, with what cleanup and what happens when it exceeds the fuse.

Streaming controls memory, not total cost

The database formats rows directly as CSV and the service compresses chunks as they arrive. This avoids constructing millions of application objects or holding the entire file in Python memory. Closing the gzip stream writes the trailer a consumer uses to detect truncation.

Bounded memory does not make the query free. Each accepted export still scans, joins, orders and transfers a large snapshot. The endpoint is therefore limited to eligible plans and throttled per key. If concurrent demand begins competing with daily ingestion or normal reads, the correct next step is global admission control, a replica or a published object — not a larger timeout.

Delivery needs evidence at both ends

An HTTP 200 only proves that a response started. It does not prove the client received a complete gzip member. A production service should observe starts, database duration, bytes emitted, cancellation and concurrency. A consumer should download to a temporary filename, validate gzip, inspect row-count drift and only then promote the new dated snapshot.

That distinction is why “just add a CSV route” is not a fair comparison. A consumer integration should fail visibly, preserve its last good snapshot and record the dated filename, observed row count and retrieval time for the dataset it loaded.

Why pay instead of running COPY yourself?

The SQL is not scarce. The maintained input and operational contract are:

If a team already operates the complete upstream history and export controls, building its own dump may be rational. Most buyers need the current dataset, not another pipeline to own. The Hosted API turns that maintenance burden into one authenticated download.

Open the company dump endpoint Compare Pro and Risk

build 2026.10.10·8073961