AI Data Bill of Materials (DBOM): Why AI Security Needs a Data Supply Chain
An SBOM tells you what code is in your software. It can't tell you which customer records ended up in your latest fine-tune. A DBOM makes the data supply chain explicit: every dataset behind a model, where it came from, how it was processed, and who owns it.
A Data Bill of Materials, or DBOM, is an inventory of every dataset that fed an AI model. It records the source of each dataset, the processing steps applied, the sensitivity level, the contractual terms governing use, and the model or agent the data ultimately shaped. If a software bill of materials answers “what code is in this binary?”, a DBOM answers the question that now opens most AI vendor reviews: what data is in this model?
Pranava Adduri, co-founder and CTO of Bedrock Data, put it in one line on episode 95: “models have a DBOM, a data bill of materials. What data made its way into the training phase? Where did that data come from?”
That definition matters because most enterprise AI programs today have no usable inventory of the data feeding their models. They have spreadsheets, lineage tools for the warehouse, and informal documentation, but nothing that survives an auditor asking which customer records ended up in the latest fine-tune. A DBOM turns that question into a lookup. (The other half of that episode, how a security leader says yes to AI without running the office of no, is in Stop Saying No: Enable AI With Data Governance.)
What is a Data Bill of Materials (DBOM)?
The term comes from Bedrock Data, whose own documentation describes a DBOM as a record that “enumerates the datasets used for AI training and inference,” creating “a verifiable record of every data source an AI model consumes” (Bedrock Data, AI Data Bill of Materials). Pranava’s framing on the show starts from something every engineer already understands:
“You have your binary that you’re vending. There’s the software dependencies of those binaries, and then there’s the downstream dependencies of those binaries as well, and libraries as well. And so in that same way, models have a DBOM, a data bill of materials. What data made its way into the training phase? Where did that data come from? What sets of processing did that data go from, original data to refined data, to get it to a stage where it’s training ready?”
Pranava Adduri, co-founder and CTO of Bedrock Data, on episode 95
Two details carry the weight. Processing is part of the record: raw data and its anonymized, training-ready version are different artifacts, and the step between them is where controls succeed or fail. And the DBOM is a snapshot of, in Pranava’s words, “the data supply chain that’s going into it,” which makes it versionable, diffable and signable: evidence rather than folklore.
You will also see the term AI-BOM for a wider inventory of models, datasets, frameworks and prompts; the CycloneDX SBOM standard already has an ML-BOM capability covering “models, datasets, and their dependencies,” including “provenance and ethical considerations for datasets.” A DBOM is the data slice of that picture, and the slice to start with, because the data is where the contractual and regulatory exposure lives.
SBOM vs. DBOM: why a software bill of materials can’t see training data
The software bill of materials has had the better part of a decade of government attention. The NTIA’s multistakeholder process began in 2018, Executive Order 14028 followed in 2021, and the resulting Minimum Elements for a Software Bill of Materials (July 2021) defined an SBOM as “a formal record containing the details and supply chain relationships of various components used in building software.” CISA’s shorthand is better still: “a nested inventory, a list of ingredients that make up software components”. Tools like Syft and Trivy generate one in a build step, and procurement asks for it almost as routinely as a SOC 2.
None of that helps when the risk lives in the data rather than the code. An AI model can be wrapped in clean, well-audited software and still expose the business if the training data contained customer PII that should never have left a regulated zone. The binary is identical whether the model was fine-tuned on synthetic data or a leaked production dump, and the SBOM can’t tell you which one happened. That is why the agentic AI security conversation now extends past the agent’s permissions and into the training set itself. The encouraging part is that the SBOM already taught your organization the habit: an inventory, a set of minimum fields, a hand-off to procurement.
If SBOMs are new to you as well, Harness (whose head of security and compliance, Andrew Spangler, was our guest on episode 26) has a five-minute explainer:
What belongs in a DBOM? Seven fields for every dataset
A useful DBOM lists more than dataset names. Borrowing the NTIA’s idea of minimum elements, here is the least to record for every dataset that touches a model:
- Source. Internal system, vendor, public corpus, customer-uploaded, synthetic. For personal data, the purpose it was originally collected for.
- License and contract terms. Granularity caps, geographic restrictions, retention limits, whether training use is permitted at all.
- Sensitivity classification. Tied to your existing scheme, re-checked after processing instead of inherited from the source.
- Processing steps. Anonymization, de-identification, sampling, filtering, labeling, with the version of the tool that did it.
- Training stage. Pretraining, fine-tuning, RAG grounding, evaluation. Data that only grounds retrieval carries a different exposure than data baked into the weights.
- Consuming model or agent. The downstream artifact that inherited this data’s risk, by name and version.
- Owner. A named person accountable for the dataset’s accuracy and compliance.
The contract terms field exists because procured data carries obligations that never show up in a schema. Pranava’s example: “In certain domains where the company might be procuring data from data vendors, there might be terms around what granularity of data can be used when training a model.” A DBOM that lists only internal datasets misses the contractual risk that comes with the purchased ones.
The owner field exists because of a distinction George Gerchow, Bedrock Data’s chief security officer, drew from his years at Sumo Logic and MongoDB. He was never the owner of the data, he told us: it “was always someone like a Pranava, a CTO type that owned it, but I was responsible for the stewardship and protection of that data. But the scale was just out of control.” Stewardship without ownership is how a dataset ends up in a training run with nobody able to say whether it was allowed to be there. A name in the DBOM fixes that at the cheapest possible point.
How does the AI data supply chain work, and where does provenance break?
Think of the data behind a model as a supply chain with five hand-offs: raw sources, processing, a training-ready set, the model or agent, and its output. Provenance is the record of each hand-off, and it breaks wherever a hand-off happens without anyone writing it down. George’s description is the one we repeat to clients: “when it comes to data provenance, from cradle to grave when data is created, but how does it move and what are the copies and does it still have the same entitlements to it?” The copies are the point. Every extract, every notebook, every “let me just grab a sample” creates a new artifact whose entitlements quietly default to whoever made it.
The fifth hand-off is the one people forget. Outputs don’t leave the system; they become inputs. Pranava calls it the data echo chamber:
“There’s the data you have, goes into AI, people ask prompts, it puts out new data, and that will get fed back in as well. As you enter into this rapid acceleration, there has to be a layer that’s understanding what’s going on at the speed of machine, not at the speed of human to keep up with it.”
Pranava Adduri, Bedrock Data, episode 95
Governments have reached the same conclusion. In May 2025 the NSA’s AI Security Center, CISA, the FBI and partner agencies in Australia, New Zealand and the UK published AI Data Security: Best Practices for Securing Data Used to Train & Operate AI Systems. The first of its ten best practices is to source reliable data and track data provenance: “Implement provenance tracking to enable the tracing of data origins, and log the path that data follows through an AI system.” It even recommends “a secure provenance database that is cryptographically signed and maintains an immutable, append-only ledger.” A DBOM is the minimum version of that ledger, and the version a Series B company can actually keep.
The failure mode a DBOM prevents: tainted training data
Most organizations, Pranava noted, have a control that says data must be anonymized before it reaches the training phase; when that step is skipped, the model is tainted and nothing in the build pipeline notices. Customer emails, account IDs or partial card numbers end up baked into the weights. The model passes evaluation and ships. Months later, someone extracts the original records with a well-crafted prompt.
George framed why a security leader cares: “What’s the first thing that pops into mind during a security incident? Did any sensitive data ever get out? Is that data being misused, mishandled?” He has lived the version where the answer arrives late: “closing the incident, Jon, and going, my goodness, did I close it? And now all of a sudden, someone’s going to put it on social media that they actually exfiltrated the data out of the environment. I just didn’t know.” When the answer means reconstructing the training pipeline from chat logs and Jira tickets, it arrives too late. When the answer is a DBOM query, it arrives in minutes, and the incident report gets to say so.
From DBOM to guardrails: how the Argus gap analysis works
A DBOM on its own is documentation. It becomes a security control when paired with runtime guardrails on the model’s output. Pranava described the pattern Bedrock built into Argus, which the company launched just before AWS re:Invent:
“When you look at a model, on the left side, compute the DBOM, what all went into it, and on the right side, look at the guardrails to figure out what do the guardrails enable, what do they block, and what do they allow. And that gives you a gap analysis based on what’s going into the model or the agent or the copilot. You have this potential of what might be coming out. What do the guardrails block? And then what do they still let through? And it’s a way of coming back to what I was originally talking about, which is responsibly enabling innovation. How do you let people run fast while monitoring the gap analysis and seeing where you might be getting in trouble?”
Pranava Adduri, Bedrock Data, episode 95
The logic fits on a whiteboard. Anything the DBOM proves went in that no guardrail blocks is a known exposure, and it goes to the top of the list. Anything a guardrail blocks that never appears in the DBOM is a rule you can stop tuning. Without a DBOM the left column is blank, and the guardrails are an act of faith rather than a control.
This is also why the DBOM is the enabling half of the story. The conversations Pranava has with CTOs start from the same place: “we roll out this AI, we want to use it to be competitive, but we don’t know what data is going in. And as a result, we can’t control necessarily what’s coming out.” Once you can answer what you have deployed and what data feeds it, the yes gets easier: you can see which models are fine and which one needs its data reworked. Or, as he put it, you get to “say, hey, you’re building fast, but this particular model might be in trouble because of the data it’s exposing.”
What do ISO 42001, NIST AI RMF and the EU AI Act expect for data provenance?
We raised ISO 42001 on the episode because it is what we hear on customer calls: a prospect’s security team asks where the training data came from and what type of data is in it, before anyone asks about the model. Pranava’s answer was that the DBOM’s contents are “elements that could be used for organizations going up for ISO audit as well. In terms of knowing what data is going in and where that data came from upstream.” Here is how the three frameworks buyers cite most often line up with a DBOM:
- ISO/IEC 42001. Published in December 2023, it “specifies requirements for establishing, implementing, maintaining, and continually improving an Artificial Intelligence Management System (AIMS)” (ISO). Its Annex A includes a control area for data in AI systems, with controls for acquisition of data (documenting sources, categories and data rights), quality of data, data provenance (a process for recording where data came from) and data preparation (documenting the methods used); Hyperproof’s walkthrough of the Annex A controls is a readable vendor summary. Each of those controls is a DBOM column. We wrote about why ISO 42001 is becoming the companion request to SOC 2 in AI Agents Are Now on Both Sides of the Breach; an ISO 42001 program is where the DBOM stops being a spreadsheet and becomes audit evidence.
- NIST AI RMF. The voluntary AI Risk Management Framework, released in January 2023 with four functions (Govern, Map, Measure, Manage) and extended by a Generative AI Profile in July 2024. A DBOM is what makes the Map and Measure work concrete for data.
- EU AI Act, Article 10. For high-risk systems, training, validation and testing datasets must meet quality criteria, with data governance covering “data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection,” and “relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation” (Article 10 text). Those are the source and processing-steps fields, almost word for word. The high-risk obligations phase in through 2027 and 2028 on the current timetable, which makes now a comfortable time to start the record.
A SOC 2, for the record, covers AI features only inside the audited system boundary and says little about training data (what’s actually in a SOC 2 report); if you handle health data, HIPAA’s de-identification rules add their own processing-step requirements. Regulators and buyers are converging on the same request: show me where the data came from and what you did to it. A DBOM answers once.
How to build a minimum viable DBOM without boiling the ocean
You don’t need a perfect DBOM on day one. You need a minimum viable one covering the highest-risk models first, and the first version takes a few weeks with a template and a calendar.
- List the models and agents in production. A real census usually finds far more AI workloads than the security team knew about, including the ones employees connected on their own. A shadow AI detection pass gives you the inventory; the case for governing those tools rather than banning them is in Shadow AI: The Risks, and How to Govern It.
- Pick the three highest-risk workloads by exposure to customer data, regulated workflows, and external-facing surface. It is the same scoping discipline that separates the AI programs that ship from the ones that stall, and the same pre-flight readiness check you would run before deploying an agent, pointed at the data.
- Document the data feeding those three first. Source, terms, sensitivity, processing, stage, consuming model, owner. Pick a template; don’t harmonize the whole company. Where a dataset is sold or shared as well as trained on, scrubbing PII and secrets and documenting rights and provenance is the same work with a price tag attached.
- Connect the DBOM to the guardrails for those three. Run the left-side, right-side comparison. Close the highest-severity items.
- Expand only after the first three are in steady state. Three models covered well beats fifty covered poorly. For new agents, write the DBOM entry at onboarding, when the agent declares the data it needs before it runs, instead of reconstructing it afterward.
Inventory, classify, monitor, remediate: the discipline is old. The novelty is that the inventoried asset is data, not endpoints. George described what it feels like once the inventory exists: “When you have this ability for data, know the lineage, the entitlements of that data, how it works, what’s sensitive, what’s not. It’s just a very powerful, comforting feeling.” It also has a commercial value: you can tell a buyer’s security team, with a straight face and a document, exactly what your model learned from.
The decision: inventory the data before the auditor asks
The decision in front of you is smaller than it looks. You don’t have to instrument every dataset in the company this quarter. You have to name your three most exposed models, write down what fed them, and compare that list with what your guardrails block. That exercise turns “we think the training data was clean” into a record you can hand to an auditor, a buyer or your own incident commander, and it is the first piece of evidence an ISO 42001 program builds on. To hear it from the team that turned the idea into a product, Pranava and George walk through the DBOM, Argus and the data-first view of security on episode 95.
Data Bill of Materials frequently asked questions
- What is a Data Bill of Materials (DBOM)?
- A DBOM is an inventory of every dataset that fed an AI model, agent or copilot. For each dataset it records the source, the license or contract terms, the sensitivity classification, the processing steps applied (anonymization, filtering, labeling), the training stage it was used in, the model that consumed it, and a named owner. It is the data equivalent of a software bill of materials: a record of what went into the thing you ship.
- What is the difference between an SBOM and a DBOM?
- An SBOM lists the software components and dependencies in a binary; a DBOM lists the datasets and processing steps behind a model. Two models can be byte-for-byte identical software while one was fine-tuned on synthetic data and the other on a production dump. Only the DBOM can tell them apart, which is why AI security programs need both.
- Is a DBOM the same as an AI-BOM?
- They overlap. AI-BOM (AI bill of materials) is the broader industry term for documenting an AI system's models, datasets, frameworks and dependencies; the CycloneDX ML-BOM format is one way to write it down. A DBOM is the data-specific slice: it goes deeper on where each dataset came from, how it was processed and what you are allowed to do with it. Most teams start with the DBOM because the data is where the regulatory and contractual exposure lives.
- Why does AI security need a DBOM?
- Because the highest-impact AI failures are data failures. A model trained on customer records that skipped anonymization can surface them through prompt extraction months later, and a model trained on vendor data outside the contract's terms is a legal problem no firewall can fix. A DBOM lets you answer "what is in this model?" from a record instead of reconstructing it from chat logs, and it is what auditors and enterprise buyers increasingly ask to see.
- What does ISO 42001 require for training data provenance?
- ISO/IEC 42001, the AI management system standard published in December 2023, includes an Annex A control area for data used in AI systems. Its controls cover acquisition of data (documenting sources, categories and data rights), data quality, data provenance (a process for recording where data came from) and data preparation (documenting the methods used). A DBOM is the artifact that satisfies those controls in one place.
- Does the EU AI Act require documenting training data?
- For high-risk AI systems, yes. Article 10 requires training, validation and testing datasets to meet quality criteria, with data governance practices covering data collection processes and the origin of data, and data-preparation operations such as annotation, labeling, cleaning and enrichment. Those are DBOM fields. The high-risk obligations phase in through 2027 and 2028 under the current timetable, so the record you start now is the one you will be asked for.
- How do you build a DBOM without a dedicated tool?
- Start small and manual. List the models and agents in production, pick the three with the most exposure to customer or regulated data, and fill in a one-page template for each dataset feeding them: source, terms, sensitivity, processing steps, stage, consuming model, owner. Store it with the model's version. Tooling such as Bedrock Data's DBOM or a CycloneDX ML-BOM can automate it later, but three models documented well is already more than most enterprises have.
- How does a DBOM connect to AI guardrails?
- The DBOM tells you what sensitive data could come out of a model; the guardrails tell you what you actively block on the way out. Comparing the two is a gap analysis: anything the DBOM proves went in that no guardrail blocks is a known exposure to close first. Without the DBOM you cannot compute that comparison, so the guardrails are configured on hope rather than evidence.