Est.

MCP Server Architecture for Structured Business Data

How to architect MCP servers for trustworthy agent access to business data.

Staff Writer · · 10 min read
Cover illustration for “MCP Server Architecture for Structured Business Data”
Agent Data Access Patterns · October 9, 2026 · 10 min read · 2,220 words

The Model Context Protocol is a standardized interface that lets AI agents interact with external systems, not a thicker layer of API glue. Anthropic launched MCP in November 2024 specifically to standardize how AI systems access external data, replacing a world where every integration between a model and a system required its own one-off code.

That separation of roles does real work. Because the server absorbs authentication tokens, rate limits, schema mapping, and error handling, the agent never has to know how the underlying system is structured. It asks for what it needs and receives an answer back in a consistent shape.

This structure also solves a scaling problem that predates MCP by years. MCP collapses that math: each new system requires one server, not one integration per model that might ever need to talk to it.

MCP's governance has caught up with its adoption. As of December 2025, MCP sits with the Linux Foundation as a founding project of the Agentic AI Foundation, a directed fund under the Linux Foundation, with Anthropic, OpenAI, and Block as founding members and AWS, Google, Microsoft, Bloomberg, and Cloudflare among the platinum sponsors alongside Anthropic, Block, and OpenAI. None of this says anything yet about whether any given MCP server returns good data. The protocol defines how an agent talks to a server. What the server does internally, and what it decides to expose, is a separate decision entirely, and it is the decision that determines whether the whole arrangement is trustworthy.

The MCP Server's Interior Design

MCP standardizes the conversation between agent and server, but it says nothing about what the server should say back. An MCP server pointed directly at raw tables does not solve the data-trust problem that has always plagued analytics. Every naming inconsistency, every permission boundary, every place where two dashboards define the same metric differently gets exposed at the speed an agent can issue queries, which is far faster than any analyst working by hand.

The distinction that matters is where the SQL gets generated and against what. An agent querying through a semantic layer retrieves governed, consistently defined metrics, because the MCP server owns the SQL generation. The same question returns the same calculation every time, because the calculation itself lives in the server, not in the agent's interpretation of raw material.

The common counterargument, that modern models are smart enough to figure out an unfamiliar data model on their own, does not hold up under production load. The model's intelligence was not the difference. It was the format of what the model was asked to reason over.

This names the real architectural question MCP adoption raises: where does the business logic live? Or it can be left to the agent, reconstructed imperfectly and differently each time a question gets asked. Those are the only two options, and the next question is what should sit in the server to make the first option work.

Pre-modeled, governed datasets as the right unit of data to expose over MCP

The right unit of data to expose through an MCP server is a pre-calculated, semantically defined dataset, not a table and not an open query path. A governed dataset encodes the business logic, the metric definitions, and the access boundaries ahead of time, so the agent consumes an answer that has already been through governance rather than raw material it has to interpret on the fly.

Not every MCP server built for business intelligence does this, and the differences between categories are not cosmetic. Pipeline-focused servers let agents build and execute data workflows. Dashboard-focused servers let agents explore visual analytics. Metrics-focused servers expose governed, semantically defined numbers directly. Choosing the wrong category for the job is one of the most common mistakes teams make when evaluating MCP tools, because all three look superficially similar from the outside: an agent asks a question, something comes back.

Getting this wrong raises latency and erodes trust. A naive MCP server forces the agent into a long chain of sequential tool calls to accomplish anything complex. Building a single dashboard with three charts pulled from different tables could require seven to ten separate tool calls under that design. A server whose interior uses agentic orchestration instead exposes one high-level tool to the agent and handles schema mapping and query generation internally, collapsing that chain into a single request.

Pre-modeling also closes off a specific failure mode familiar to anyone who has sat through a meeting where two departments argue over whose dashboard is right: MRR meaning one thing in finance's view and something else in sales'. When the metric definition lives inside the server rather than in each consumer's separate interpretation of the underlying tables, every consumer, whether a human opening a dashboard or an agent answering a question, gets the same number from the same calculation. Governance, in this architecture, is a property of where the calculation physically happens.

What the Supabase MCP Server Does for Analytics Workloads

Supabase's MCP server gives an agent broad, genuinely useful access to a Supabase project: querying data, listing tables, managing migrations, reading logs, and running security and performance advisors. Pointing it at a live production project is the wrong default, by Supabase's own account of its purpose.

Three hardening steps form the mandatory baseline for anyone running it anyway: point the server at a development project rather than production, run it with the --read-only flag so queries execute as a read-only Postgres user, and scope it with --project-ref to bound the blast radius to a single project while disabling account-level tools.

Read-only access is necessary, but it does not close every risk on its own. A model with read access to private data remains vulnerable to prompt injection, and an injected prompt can still cause that data to leak through the model's own response text even when no write-back path exists. Read-only protection closes the write-path risk. It does nothing for the read-path risk.

For analytics workloads specifically, a further precaution helps: connecting the PostgreSQL MCP server to a read replica rather than the primary database keeps AI-driven queries from competing with production traffic for resources. That is a sound operational practice, but it does not change what the agent is looking at. It is still raw, unmodeled tables, replica or not.

That is the gap. Hardening the Supabase MCP server protects production from the agent. It gives the agent access to raw schema instead of governed, pre-calculated metrics, and that pushes the entire burden of correct metric calculation onto the model's SQL generation at the moment the query runs, which is precisely the failure mode the earlier sections describe. A development tool, hardened correctly, is still a development tool. Analytics workloads need something standing between the agent and the raw database.

The architecture that puts governed data between the agent and Postgres

For teams running on Postgres, the production-safe pattern is a layered system in which pre-modeled, governed datasets sit between the agent and the live database, with the MCP server permitted to see only those datasets, rather than a single MCP server pointed at a database however carefully scoped.

The first layer is the production database itself: the Postgres primary, protected entirely from agent access. No MCP server touches it directly, under any configuration.

The second layer is the governed data layer: pre-calculated metrics and modeled datasets stored in a format agents can query cheaply and quickly. Business logic, metric definitions, and access boundaries are encoded here, at the point where the data is modeled, rather than delegated to the agent to reconstruct at query time.

The third layer is the MCP server itself, a stateless service that exposes only the governed datasets from layer two. Agents query this layer. They never reach layer one.

At production scale, the MCP server should run as an independent, stateless HTTP service that interacts with Supabase only for context persistence and authentication, using Supabase Realtime for event-driven context updates and Edge Functions for lightweight MCP endpoints. That topology keeps the server horizontally scalable while Postgres row-level security handles tenant isolation underneath it, so the server's own statelessness does not become a bottleneck as usage grows.

This is not a pattern unique to smaller Postgres-native teams improvising around a constraint. Salesforce's Headless Data 360 for MCP applies the same architectural principle at enterprise scale, delivering governed customer context directly to agents so that any MCP-compatible client can dynamically discover and invoke the data it needs without being told exact field names or function calls in advance. The specifics differ, but the principle holds across both: governance belongs in a layer the agent cannot bypass, not in a layer the agent is trusted to navigate correctly on its own.

What this architecture buys beyond protection from accidental damage to production is consistency. The same metric query returns the same result regardless of which agent asks, which model version is running, or which user triggered the request, because the definition lives in the governed layer rather than in whatever the prompt happened to say that day. That consistency is also what makes the output auditable after the fact, which matters as much to a finance team reconciling numbers as it does to an engineer debugging a model's behavior.

Defining Metrics for the Governed Layer

A governed data layer is only as trustworthy as the metric definitions encoded inside it. No architecture, however well designed, can substitute for the upstream decision about what a given number actually means.

Most single-source-of-truth failures are organizational failures, not tooling failures. If MRR means booked subscriptions on one dashboard and recognized revenue on another, routing both through an MCP server does not resolve the disagreement. It delivers the same broken disagreement to agents faster than before, because the company never agreed on the definition. Fixing that requires a decision, not a deployment.

The specific metrics that belong in a governed layer will vary by business, but the pattern for SaaS companies is reasonably settled: MRR, ARR, Net Revenue Retention, Logo Churn, CAC Payback Period, LTV/CAC ratio, and Burn Multiple are common candidates, and each one requires a single agreed-upon definition before anyone encodes it into the layer.

A number by itself carries little information. Each metric in the governed layer should carry three values: the current figure, the trend showing change against the prior period, and a benchmark representing a target or peer comparison. Trend and benchmark together convert a raw figure into a signal that both human readers and agents can actually interpret and act on.

The metrics worth governing have also shifted as AI has moved into daily workflows. Three categories that were rare two years earlier had become standard by 2026: AI-assisted throughput, comparing work completed with AI assistance against work completed without it on measures of quality and rework; cost per outcome, inclusive of model spend; and data freshness, tracked as a first-class KPI. Teams deploying agents should build these into the governed layer from the start.

The governance point here is simple and easy to underestimate. Automating metric delivery through an MCP layer accelerates the distribution of whatever definition is encoded in it, for better or worse. Governance has to come before the feature does. The definition needs to be fixed before the agent starts reporting it, because once an agent is reporting a number on a schedule, correcting a bad definition means correcting every report that number has already touched.

Agent-to-Data Access in Production

When the governed layer is actually in place, agents stop functioning as unreliable fetchers of raw data and start functioning as reliable analytical collaborators. That shift comes from the architecture surrounding the model, not from any improvement in the model's own reasoning ability.

Ramp's production case shows the pattern directly. Claude then generates SQL against that structured, tabular data rather than attempting arithmetic on raw JSON responses, and the result was a model that went from struggling with a few hundred data points to accurately analyzing tens of thousands of spend events.

The model is considerably better at generating SQL against clean, tabular data than it is at performing arithmetic over unstructured JSON, and that single observation justifies the entire architectural investment in a governed layer.

With that layer in place, a new kind of automated workflow becomes possible. An agent running on a schedule can read the latest data snapshots from the governed layer, identify which metrics have moved beyond a set threshold, and produce a first draft explanation of what changed and why. The human decision-maker on the receiving end gets analysis to review, not a raw data-fetch task to perform manually before the real thinking can start.

This fails if the layer underneath is wrong. Treating the agent as the source of truth rather than the governed layer beneath it is the failure mode to guard against above all others: automating a broken metric definition through an agent only distributes the broken definition faster and with more apparent authority than before. The architecture earns its keep when the governed layer it depends on is correct.

Freshness deserves a place in that layer as a first-class metric in its own right. Agents need to know when a number was last calculated, because stale context delivered with total confidence carries its own category of risk, and a well-designed architecture should surface that risk.

More in Agent Data Access Patterns