{"id":5243,"date":"2026-10-07T10:50:04","date_gmt":"2026-10-07T10:50:04","guid":{"rendered":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/07\/federated-query-at-petabyte-scale-a-deployment-pattern-for-a-governed-ai-agent-data-layer\/"},"modified":"2026-10-07T10:50:04","modified_gmt":"2026-10-07T10:50:04","slug":"federated-query-at-petabyte-scale-a-deployment-pattern-for-a-governed-ai-agent-data-layer","status":"publish","type":"post","link":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/2026\/10\/07\/federated-query-at-petabyte-scale-a-deployment-pattern-for-a-governed-ai-agent-data-layer\/","title":{"rendered":"Federated Query at Petabyte Scale: A Deployment Pattern for a Governed AI-Agent Data Layer"},"content":{"rendered":"<div><img data-opt-id=1285383124  fetchpriority=\"high\" decoding=\"async\" width=\"770\" height=\"300\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/federated-query-genai-770x300-1.jpg\" class=\"attachment-large size-large wp-post-image\" alt=\"\" \/><\/div>\n<p><img data-opt-id=166565596  fetchpriority=\"high\" decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/devops.com\/wp-content\/uploads\/2026\/10\/federated-query-genai-770x300-1-150x150.jpg\" class=\"attachment-thumbnail size-thumbnail wp-post-image\" alt=\"\" \/><\/p>\n<p><span>Enterprises with data spread across ERP, CRM, cloud warehouses and observability platforms face a recurring problem: As data volume grows into the tens of petabytes, traditional ETL-and-centralize approaches produce unacceptable query latency and duplicated infrastructure cost. This case study describes an automated deployment framework for a distributed SQL query engine (federated query architecture) that queries data at its source <\/span><span>rather than moving it, built and operated<\/span><span> as the lead infrastructure engineer at a health care technology company. It further describes how that same data layer was extended into a governed, agent-accessible interface for enterprise GenAI, using a semantic layer and tool-exposure protocol (MCP-style) to connect a conversational AI interface to real-time business data. This architecture fully preserves the role-based access control and audit requirements appropriate for regulated (health care-adjacent) data. The contribution is a reusable deployment and governance pattern, not a speci\ufb01c vendor product.<\/span><\/p>\n<h3><span>Problem<\/span><\/h3>\n<p><span>A business intelligence function at a health care technology company had grown to ingest data from multiple ERP\/CRM systems, multiple cloud data warehouses and observability\/logging platforms, converging on roughly 30 petabytes of federated data. The broader data and analytics organization supporting this initiative numbered 30\u201335 people; the automated deployment framework and federated query infrastructure speci\ufb01cally were owned and built by <\/span><span>this author<\/span><span> as the team\u2019s DevOps engineer. This growth happened faster than the data architecture could absorb: Each new source system arrived with its own connector requirements, its own access model and its own idea of what \u2018current\u2019 data meant.<\/span><\/p>\n<p><span>The default response to this kind of sprawl is to centralize \u2014 build a pipeline that copies everything into one warehouse, then query the copy. At this scale, that approach broke down in a few speci\ufb01c ways:<\/span><\/p>\n<ul>\n<li><span>Query turnaround had no consistent \ufb02oor. Depending on which systems a business question touched, getting an answer could mean waiting on a batch ETL cycle, requesting a manual data pull from another team or stitching together exports by hand. There was no reliable answer to how long it would take. This unpredictability was often more damaging to the business than any single slow query.<\/span><\/li>\n<li><span>Storage and pipeline maintenance were duplicated. Centralizing meant maintaining a second (or third) copy of data that already had an owner and a home. Every schema change upstream meant a corresponding \ufb01x downstream, multiplied across every consuming pipeline.<\/span><\/li>\n<li><span>Governance got \ufb02attened. Once data left its source system and landed in a shared warehouse, the \ufb01ne-grained access controls that existed at the source often didn\u2019t carry over cleanly \u2014 meaning either over-provisioning access or building a second permissions system from scratch.<\/span><\/li>\n<\/ul>\n<p><span>The alternative \u2014 federated query, where the engine queries data in-place at the source rather than copying it \u2014 solves the duplication and staleness problems but shifts the burden onto operating the query engine itself reliably. When this work began, distributed SQL engines built for this pattern (Starburst\/Trino-based architectures) were still a relatively young category, and the available documentation focused on manual, one-time cluster setup. There was very little guidance on making that deployment repeatable, automated and safe to redeploy without losing operational history, which is the gap this work addressed directly.<\/span><\/p>\n<h3><span>Approach: Automated Federated Query Deployment<\/span><\/h3>\n<p><span>The core contribution was an integrated CI\/CD deployment framework for a Kubernetes-hosted, distributed SQL query engine, built from the ground up rather than adapted from an existing internal playbook, as no such playbook existed for this engine at this scale.<\/span><\/p>\n<ul>\n<li><span>Automated Cluster Provisioning and Upgrade Pipelines: Cluster setup, con\ufb01guration and version upgrades were previously handled manually \u2014 a process that took 3<\/span><span>\u2013<\/span><span>5 days per deployment cycle when done by hand and carried real risk of con\ufb01guration drift between environments (a setting changed in staging but forgotten in production, for instance). Moving this into a versioned CI\/CD pipeline meant every deployment used the same tested con\ufb01guration, and the same pipeline that deployed to a test environment could promote to production with a predictable, auditable set of steps. This took deployment time down to a few hours \u2014 roughly a 10x reduction \u2014 but the more durable bene\ufb01t was consistency: Every engineer on the team could trigger a deployment with the same expected outcome, rather than the process depending on one person\u2019s institutional knowledge of the manual steps.<\/span><\/li>\n<li><span>Log Persistence Across Deployments: Kubernetes pods are ephemeral by design. When a pod restarts or is redeployed, anything written only to its local \ufb01lesystem is gone. For a query engine handling business-critical data, losing operational logs on every redeploy meant losing the ability to diagnose what happened before an incident. The available documentation for this engine, being relatively new to the ecosystem, didn\u2019t have a clear answer for this. The solution was to route logs to persistent, external storage independent of pod life cycle, so redeploys (routine or emergency) never cost the team its operational history.<\/span><\/li>\n<li><span>Connector-Based Access to Source Systems: Rather than extracting data from ERP, CRM, cloud object storage and cloud data warehouse systems into a central store, the query engine connects to each source directly and executes queries against the live data. This eliminates the staleness problem inherent to batch ETL (data is only ever as old as the source system, not as old as the last successful pipeline run) and eliminates the duplicated storage cost of keeping a second copy of everything.<\/span><\/li>\n<li><span>Least-Privilege IAM as Code: Access control was provisioned through Terraform rather than con\ufb01gured manually, so that permissions granted to the federated query layer mirrored the access controls already in place at each source system, rather than introducing a new, separately managed permission model. This mattered speci\ufb01cally because a portion of the underlying data is subject to health care data protection requirements (HIPAA) \u2014 access needed to remain auditable and scoped, not \ufb02attened into a convenient but overly broad tier just because the data now sat behind a di\ufb00erent query interface.<\/span><\/li>\n<li><span>Outcome: Business users can now run a single SQL query joining data across previously siloed systems and get a result in 15\u201330 minutes \u2014 down from a process that, before, had no consistent turnaround time at all, since it depended on ad hoc manual work across teams.<\/span><\/li>\n<\/ul>\n<h3><span>Extension: From Query Layer to AI-Agent Data Layer<\/span><\/h3>\n<p><span>Once the federated query layer was stable and fast, the natural next request from the business was: Could people just ask questions in plain language instead of writing SQL? That request is where most enterprise GenAI projects get into trouble. The easy version is to point an LLM at a document store or a database and let it retrieve whatever seems relevant, and this version is exactly what breaks down when some of that data is regulated.<\/span><\/p>\n<p><span>The pattern built here keeps the LLM several layers <\/span><span>removed<\/span><span> from raw data at all times:<\/span><\/p>\n<ol>\n<li><span>Data Products Layer: Rather than exposing raw tables, the federated query engine surfaces curated, named datasets \u2014 volume, feedback\/case data, subscriptions, territory data and similar business-de\ufb01ned products \u2014 as stable, documented interfaces. This layer is the \ufb01rst checkpoint: Nothing downstream ever sees a raw table, only a data product someone has deliberately de\ufb01ned and named.<\/span><\/li>\n<li><span>Semantic\/Governance Layer: A knowledge graph and catalog layer sits on top of the data products, carrying context (what does this \ufb01eld mean, where did it come from), lineage (what upstream sources feed it) and access governance (who is allowed to see it). This is what lets an AI agent\u2019s access be scoped and auditable, rather than \u2018the model can see whatever the underlying database permissions allow\u2019 \u2014 which is rarely a policy anyone actually designed on purpose.<\/span><\/li>\n<li><span>Tool-Exposure Layer: A protocol-based server (following the emerging MCP pattern) exposes the governed data products as discrete, callable \u2018tools\u2019 that an LLM-based orchestrator can invoke \u2014 with de\ufb01ned inputs and outputs \u2014 rather than granting the LLM a live database connection. The distinction matters: A tool call can be logged, rate-limited and scoped to a speci\ufb01c data product. A database connection generally can\u2019t be constrained the same way once the model is composing its own queries.<\/span><\/li>\n<li><span>Orchestration Layer: A commercial conversational AI product consumes the tool-exposed data to answer natural-language business questions inside a collaboration platform used company-wide. SQL generated in response to a user\u2019s question is constrained to the governed data products de\ufb01ned in step 1, and governing policies are applied before an answer reaches the end user \u2014 so the system\u2019s freedom to interpret a question doesn\u2019t extend to freedom to reach data it wasn\u2019t given access to.<\/span><\/li>\n<li><span>Compliance Boundary: Since a portion of the underlying data is subject to HIPAA, encryption, role-based access and audit logging are enforced at the data-product and semantic layers speci\ufb01cally \u2014 meaning the AI orchestration layer never receives ungoverned access to raw regulated data, regardless of how a user phrases their question.<\/span><\/li>\n<\/ol>\n<p><span>The result \u2014 now in use by an estimated 300<\/span><span>\u2013<\/span><span>500 business users across North America, LATAM, EMEA and APAC \u2014 is a system where a regional sales lead can ask a plain-language question in a chat interface and get an answer grounded in live, governed data without that convenience requiring anyone to loosen the access controls that existed before the AI layer was added.<\/span><\/p>\n<p><span>This is, in e\ufb00ect, a retrieval-and-tool-use pattern for enterprise AI agents that preserves existing data governance, rather than a RAG pipeline built on an unstructured document dump.<\/span><\/p>\n<h3><span>Why This Matters Beyond one Company<\/span><\/h3>\n<p><span>The pattern addressed here \u2014 automate federated query deployment \ufb01rst, then expose that same governed layer to AI agents via a semantic\/tool layer rather than piping raw data into an LLM \u2014 is broadly applicable to any enterprise trying to make GenAI safe to use against regulated or sensitive data. The speci\ufb01c components (query engine choice, LLM vendor, protocol implementation) are replaceable; the sequencing and governance boundary are the reusable insight:<\/span><\/p>\n<ul>\n<li><span>Don\u2019t let the AI orchestration layer talk directly to source systems or raw warehouses.<\/span><\/li>\n<li><span>Put a data-products layer and a semantic\/governance layer between the two.<\/span><\/li>\n<li><span>Automate the deployment of the query layer itself, or it becomes the bottleneck that makes \u2018real-time AI answers\u2019 impossible at petabyte scale.<\/span><\/li>\n<\/ul>\n<h3><span>Results<\/span><\/h3>\n<ul>\n<li><span>Deployment Time: Manual Starburst cluster deployment\/upgrade previously took 3\u20135 days. The automated CI\/CD pipeline reduced this to a few hours \u2014 a 10x reduction.<\/span><\/li>\n<li><span>Query Turnaround: Cross-source business queries that previously had no consistent turnaround time (dependent on ad hoc extracts and manual joins across systems) now typically return in 15\u201330 minutes through the federated query layer.<\/span><\/li>\n<li><span>Data Volume Under Management: It is approximately 30 petabytes, federated across ERP, CRM, cloud data warehouse and observability sources.<\/span><\/li>\n<li><span>Adoption: The natural-language AI interface built on top of this data layer is in use by an estimated 300\u2013500 business users across North America, LATAM, EMEA and APAC.<\/span><\/li>\n<li><span>Delivery: The initial version of the automated deployment framework was designed and released over 2\u20133 months, supporting a broader data and analytics organization of 30\u201335 people, with the deployment automation and query-layer work itself owned by <\/span><span>this author<\/span><span>.<\/span><\/li>\n<\/ul>\n<h3><span>Conclusion<\/span><\/h3>\n<p><span>Enterprises adopting GenAI over their own data face a governance problem before they face a modeling problem. This case study describes a deployment and architecture pattern \u2014 automated federated query infrastructure plus a semantic\/tool-exposure layer \u2014 that let one infrastructure engineer scale a petabyte-class data platform into a governed, AI-agent-accessible system without compromising regulated-data controls. <\/span><\/p>\n<h3><span>Key Takeaways<\/span><\/h3>\n<ul>\n<li>\n<ul>\n<li><span>At petabyte scale, federated query (query-in-place) avoids the staleness, duplicated storage and \ufb02attened governance that centralize-then-query approaches introduce.<\/span><\/li>\n<li><span>Automating the deployment of the query engine itself, not just the data pipelines, is what makes federated query operationally viable cutting a 3\u20135-day manual process to a few hours.<\/span><\/li>\n<li><span>Ephemeral infrastructure such as Kubernetes pods requires deliberate log persistence design, or operational history is lost on every redeploy.<\/span><\/li>\n<li><span>Making enterprise data safely available to an LLM agent requires a data-products and semantic\/governance layer between the model and raw sources, not direct database access.<\/span><\/li>\n<\/ul>\n<\/li>\n<li><span>A tool-exposure layer (MCP-style) lets an AI agent\u2019s access to governed data be logged, rate-limited and scoped in ways a raw database connection cannot be \u2014 preserving existing compliance controls such as HIPAA even as AI is added on top. <\/span><\/li>\n<\/ul>\n<p><a href=\"https:\/\/devops.com\/federated-query-at-petabyte-scale-a-deployment-pattern-for-a-governed-ai-agent-data-layer\/\" target=\"_blank\" class=\"feedzy-rss-link-icon\">Read More<\/a><\/p>\n<p>\u200b<\/p>","protected":false},"excerpt":{"rendered":"<p>Enterprises with data spread across ERP, CRM, cloud warehouses and observability platforms face a recurring problem: As data volume grows [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":5244,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","ast-disable-related-posts":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"default","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"var(--ast-global-color-4)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":""},"categories":[5],"tags":[],"class_list":["post-5243","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-devops"],"_links":{"self":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5243","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/comments?post=5243"}],"version-history":[{"count":0,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/posts\/5243\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media\/5244"}],"wp:attachment":[{"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/media?parent=5243"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/categories?post=5243"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/rssfeedtelegrambot.bnaya.co.il\/index.php\/wp-json\/wp\/v2\/tags?post=5243"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}