Read Replicas vs Governed Datasets for Agent Queries
Governed datasets embed security and semantics that read replicas cannot provide for AI agents.

Read replicas solve a specific, well-understood problem: they keep heavy analytical load away from the transactional database that an application depends on for predictable latency. A replica is a copy of production, updated continuously through replication, sitting on its own compute so that a slow report or a wide scan does not compete with the writes and reads an application needs to stay fast. The appeal of this pattern runs deep in any team that already lives in Postgres, because the interface on the other side is identical Postgres: same SQL dialect, same driver, same mental model. A data analyst who already knows how to write a join against the production schema can point that same query at the replica and get the same answer, just without putting the primary at risk.
For a human analyst running an occasional cohort query or compiling a weekly report, that arrangement works well. The analyst brings judgment to the table: knowledge of which columns actually hold the number finance cares about, awareness of which tables are stale or deprecated, and the patience to wait for a result and sanity-check it before using it. The replica's job was never to carry that judgment. Its guarantee is narrow and specific: reads don't hit production. That promise says nothing about whether the query itself is safe, whether the result is semantically correct, whether the requester should have seen that data in the first place, or whether anyone could later reconstruct what was queried and why. Those gaps were tolerable when the only consumer on the other end was a credentialed person who understood the data model well enough to catch their own mistakes.
How agent queries differ from human analyst queries
An AI agent is not a faster analyst but a structurally different kind of data consumer. Agents run continuously, at machine speed, without the natural pauses that come from a person logging off at the end of the day. A replica provisioned with broad read access in January may still be executing those same permissions months later, long after the business purpose that justified the access has changed or disappeared entirely, because nothing about the architecture forces a review.
Agents also lack the semantic context a human analyst brings. Without governance embedded directly into the data they touch, an agent has no way to distinguish between a table it was meant to use and a table it merely has the technical permission to reach. Given a choice between two columns that both look like "revenue," it will pick one, run with it, and pass the result downstream without flagging any uncertainty. This is precisely why platforms like Dreambase pre-model datasets into governed, business-meaning definitions: instead of an agent querying raw operational tables and guessing which column represents the right metric, it accesses pre-calculated, centrally defined measures that embed the semantic correctness the agent cannot infer on its own.
The problem compounds when agents start calling other agents. A compliance agent operating entirely within its stated permissions can still end up ingesting unmasked personal data if it receives that data as an intermediate output from a different agent's workflow. No single agent exceeded its access boundary, and the violation happened anyway, in the handoff between them. A human analyst who got confused about a column would ask a colleague. An agent picks a table and acts. None of this is something a replica was ever built to catch. A replica routes a query to different hardware and stops there; it adds nothing to an agent's ability to reason about what it is looking at, and it removes nothing from the risk that the agent reasons wrong.
The three specific ways read replicas fail under agent traffic
Routing agent traffic to a read replica creates three failure modes, visible together because each stems from the same mismatch between replica design and continuous machine-speed querying.
The first is a correctness failure rooted in how replication actually works. Postgres can cancel a long-running query on a replica when incoming replication data conflicts with the rows that query is reading. A multi-minute analytical scan, the kind of workload agents generate when pulling context for a report or a decision, can be terminated mid-run the moment the primary updates a row the query touches. The agent does not get a clean failure it can retry sensibly. It gets an error, or worse, a partial result, and it may act on that partial result as though it were complete.
The second is a cost structure that punishes exactly the traffic pattern agents produce. A replica runs at the same compute size as the primary and bills as a separate instance, a cost structure built around the assumption of steady, occasional, long-running human queries. Agents don't behave that way. They generate short, frequent, spiky bursts of queries, the opposite of what makes replica compute economical, and warehouse-style billing minimums common in this space compound the cost of serving that pattern.
The third is a governance surface problem, and it is the one most likely to go unnoticed until something goes wrong. A replica still exposes the same ungoverned governance surface as the primary: same schema, same column names, same broad access by default. When an agent authenticates through a shared service account, that account's combined privileges become the agent's privileges, and the end user's actual permissions disappear from the picture. Lineage is harder to maintain on a copy of the data than on the source. Freshness becomes a guess because nothing confirms the copy matches the current state of the source. Governed datasets solve this by sitting between the agent and production: agents query pre-calculated Parquet datasets with embedded access boundaries rather than raw tables, so lineage stays auditable, freshness stays explicit, and the replica never has to be governed as its own separate risk surface, on top of the one the primary already represents.
What "governed dataset" means in architectural terms
A governed dataset is a different contract between the data layer and whatever consumes it, built around federated, governed, contextual data access for agentic AI in production.
The data itself is pre-modeled before an agent ever touches it. Metrics get calculated ahead of time, definitions live in one centralized place, and the shape of the dataset reflects what the business actually means by a term rather than the raw structure of the operational schema. An agent asking for monthly recurring revenue gets the governed calculation directly, not a table of invoice rows it has to interpret and potentially misread.
Access control works differently too. Attribute-based access control enforces rules at query time rather than at login time: row-level security returns only the records an agent is authorized to see, and column-level masking redacts sensitive fields even within a query the agent is otherwise permitted to run. A fraud detection agent can see transaction amounts and timestamps while never seeing the payment card number attached to them, because the governance layer filters what comes back before the agent ever sees it.
A governed dataset can record not just what an agent queried, but the reasoning path that led to a particular query over the alternatives available to it, closer to a decision artifact than a simple access log. That distinction starts to matter once compliance obligations need to be enforced before a query runs rather than audited after the fact, with datasets themselves registering what they can be used for and under what conditions, instead of leaving that enforcement entirely to whoever wrote the access policy.
MCP as the delivery layer for keeping agents off the production database and its replicas
An open standard for connecting AI models and agents to external tools and data sources through one consistent interface is becoming the common way agents reach enterprise data platforms. It is becoming the common connective layer between agents and enterprise data platforms, and that adoption is genuine progress. But MCP solves a delivery problem, not a trust problem. The standard moves data and actions to the agent efficiently; it says nothing about whether the data on the other end is correct, current, or something the agent should have been allowed to see. An agent connected through MCP to tables that were never certified for agent use will still produce confident, wrong answers, just delivered through a cleaner pipe.
The governance has to live inside the access layer the agent actually queries through, not in a policy document sitting next to it. A zero-trust query path authenticates the caller, authorizes the specific metric being requested, logs the SQL that ran, and inspects what leaves the system on egress, all without trusting the agent's own prompt text to self-limit what joins it attempts. Governance that exists only as written policy, with no enforcement in the actual path a query travels, governs nothing in practice.
Read, export, and write paths need to be kept structurally separate. Read access granted to an agent should not include the ability to trigger a bulk CSV export from the agent's own interface, should not include an off-hours full-table extract, and should never include a write, delete, or update against the production record. A read replica still exposes the same ungoverned surface as the primary it copies, same schema, same columns, same default-broad access, and once an agent authenticates through a shared service account, the real permissions of whoever is behind that agent vanish from view. Governed datasets close that gap by sitting between the agent and production, so the production database, replica included, never has to be in the agent's query path.
Minimum viable implementation for a governed dataset layer
A governed dataset layer that agents can actually be trusted to use requires four things working together. Removing any one of them reverts the architecture to an ungoverned query path wearing governed language.
The first is pre-modeled metrics with centralized definitions: one calculation logic per metric, shared by every consumer that asks for it. The metric layer answers the question "what is MRR" before an agent ever asks it, rather than letting the agent interpret raw transaction tables and arrive at a number that disagrees with what finance already reported.
The second is identity-aware access bound at query compile time, not at login. The agent's role gets bound when the query compiles, and it gets only the connectors and metrics its specific task requires. A standing warehouse-admin credential sitting on an agent's service account fails this test immediately, no matter how well-intentioned the original setup was.
The third is a per-query audit log with full replay capability. That log needs to carry the policy version in effect at the time, the exact SQL that ran, and the result of the egress check, because without all three pieces, what happened was that a query got allowed, not that it got governed. This is the line that separates a genuinely governed dataset from a read replica that simply has row-level security switched on.
The fourth is an exposure layer agents can reach without ever touching production: pre-calculated datasets stored in an open columnar format, queryable through DuckDB or an equivalent engine. Traditional detective controls, the monthly access review, the quarterly audit, were built for human-speed violations, and they cannot catch an incident that unfolds in milliseconds. Enforcement in this layer happens at runtime, as each query executes.
Implementing the architecture for Supabase and Postgres teams
Dreambase is built to sit as the governed dataset layer between a team's Postgres data and any agent or human consumer that needs it, without requiring a data warehouse, an ETL pipeline, or a dedicated data team to maintain either one. It pre-models data pulled from Supabase, Postgres, and other connected sources into governed Parquet datasets that are queryable through DuckDB, so agents work against pre-calculated metrics rather than raw production tables and the OLTP database never enters the agent's query path.
Everything Dreambase exposes runs through a single MCP server, giving agents fast, accurate, cheap-to-query context without ever handing them credentials to the production database or to a replica sitting behind it. A governed dataset inverts the usual contract: rather than giving an agent direct access to Postgres and hoping it stays within bounds, the data it actually needs, metrics, cohorts, flagged records, gets pre-calculated and pre-shaped ahead of time, served as immutable datasets with governance built into the schema itself rather than bolted on through permissions after the fact. That decoupling is what lets agents get fast, trustworthy context without ever touching production directly.
Metric definitions stay centralized and shared across the whole system: one definition of MRR, one definition of active users, the same number whether it shows up in an agent's answer, a dashboard, or a board report, which structurally prevents the familiar disagreement between finance's number and product's number. None of this requires new infrastructure, a schema migration, or an ETL pipeline. Dreambase extends an existing Supabase project into full analytics capability from where the data already sits, which matters most for founders and operators who don't have a data team on staff but still need board-ready numbers and agents that can be trusted to work from the same ones.
When a read replica is still the right tool
None of this makes a read replica the wrong architecture in every case. A replica remains the right tool when a human analyst needs ad-hoc SQL access, when query volume stays low and predictable, when the analyst brings their own semantic judgment to the query, and when the team is willing to pay for a second compute instance sized the same as the primary. That is precisely the use case the pattern was built for, and it still serves it well.
A replica stops being the right tool once agents are the ones querying continuously at machine speed, once the organization needs a reliable record of what was queried and why, once metric definitions have to stay consistent across every consumer touching them, or once the query load turns spiky and short-duration, which is exactly the pattern replica billing models penalize hardest. A team that can no longer justify governing a second full copy of production as its own separate risk surface, on top of the one the primary already represents, should treat a replica as the wrong tool. Architectures that query data in place, without copying or moving it, keep lineage chains intact and every query traceable back to its source, with no shadow copies left to audit later.
For a Supabase team that already runs a replica for analytics, the practical path does not require tearing anything out. The replica can keep serving ad-hoc human queries exactly as it does today. A governed dataset layer takes over agent traffic and automated reporting, so each kind of consumer runs on the infrastructure actually built for how it queries, rather than forcing both onto one pattern that was only ever built for one of them.
Sources
- Architecture overview - Model Context Protocol
- Read replicas in Azure Database for PostgreSQL Flexible Server - Azure Database for PostgreSQL
- Working with read replicas for Amazon RDS for PostgreSQL - Amazon Relational Database Service
- Agentic Data Environments
- Securing Agentic AI: From Per-Action Checks to Trajectory Assurance
- The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
- A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
- Autonomous AI agents acting safely within governed data-sharing environments - International Data Spaces

