How Do You Build an AI System Inventory for Compliance?
Learn how to build an AI system inventory for compliance - what fields to capture, how to find hidden AI systems, and how to keep it up to date.

Rushali
Last quarter, a 400-person fintech got a two-week heads-up that a large customer demanded assurance that all its AI systems would comply with its data before renewing a contract. The company had a tidy record of its underwriting model. It didn't have the fraud-scoring add-on that it had agreed to purchase eight months ago, the transcription bot linked to Zoom, or the support agent to whom an engineer connected over a weekend via Slack. Most AI governance programs fall short at that point. The actual question, one that all frameworks presume you've already answered, is: how do you construct an AI system inventory for compliance without classification or documentation? This will explain what to look for, how to discover systems not logged by anyone, and how to categorize what you find, as well as where agentic AI and MCP servers fit into that.
Why You Can't Classify What You Haven't Inventoried
In almost all compliance guides to the EU AI Act, you will start the document by assessing your systems according to risk levels, see if you are a provider or a deployer, and create the necessary documentation. In the middle of that process, compliance consultancy Aguardic has identified a problem: You cannot assign a tier or a role to a system that doesn't exist. It works with each duty alone. Technical documentation required under Article 11 must contain a description of the system's design, training data, and the purpose it is to be used for; therefore, it is necessary to have a name for the system. The performance and incident tracking to meet post-market monitoring requirements (under Article 72) cannot be achieved in a system that has never been tracked. Registration in the EU database, pursuant to Article 71, requires the same starting point.
This is achieved structurally in NIST's AI Risk Management Framework. The Map function is the first of an organization's four functions in which an organization defines the context of a system, its data, its stakeholders, and its intended use. No one has captured the reality of implementing the framework more succinctly than NeuralTrust's implementation guide to the framework; Map is the AI system inventory, if it's done correctly; and it's what the artifact auditors and regulators review. If you skip it, then there's no measurable operation to run Measure and Manage. The inventory is not just a piece of paper that accompanies actual compliance effort. It's the foundation on which everything else is created.
What an AI System Inventory Actually Needs to Capture
Many organizations already maintain an inventory of the AI models and versions deployed and confuse this with the whole picture. The field list that shows only what the model does will pass that first skim and will not pass an actual audit. The following are more representative of an actual AI asset inventory for compliance: the fields that regulatory frameworks and the incidents that reveal their lack consistently refer to.

Core Identification Fields
Each entry must have a unique system ID, a plain name, a vendor in case it's third-party, a version or build in production, and a date when it was deployed. This may seem simple, but with three teams using the same fraud model and each calling it by a different name, it can be difficult to get a handle on this. What is really important about a consistent ID format is that it is what every downstream field (risk tier, owner, evidence links) is attached to. Pay attention to the deployment type as well: built-in, bought SaaS, embedded vendor feature, or API-based service. They each have different discovery paths, which are described in the next section.
Function and Purpose
Include one or two sentences, written in plain English, that give readers an idea of what the system really does, rather than what the vendor's marketing page claims it does. Giving a viewer the words “scores” and “loan applications for default risk” tells a lot more than “AI-powered credit intelligence platform.” A legal/compliance reviewer (with zero ML expertise) is the one who relies on this field to make a first-pass risk call, and it must be robust enough to “pass” out of engineering's jargon.
Connected Tools and Action Surface
It is here that most templates end up when AI systems transition from generating text for human consumption to performing actions on their own. A support ticket triage system that just provides a summary for an agent to read has a different risk profile than one connected to a refund API, customer database, and a Slack channel it can post to without them having to ask a question. Writing the model without writing the content to which it is connected says what the model can say but not what it can do. For each of the systems, identify the tools, APIs, and data stores the system can utilize, and if it utilizes them with or without the approval of a human. This field also represents the natural road map into agentic AI and MCP-connected systems, which are given a whole new section further down, as they require more than one line item.
Regulatory Fields
Record the EU AI Act risk category (unacceptable, high, limited, or minimal) and the Annex III risk category if it is a high-risk system; record the organization's role in the system (provider, deployer, importer, or distributor); record the jurisdiction(s) where the system is used. The same hiring-screening tool implemented in five EU member states and Colorado has a wider compliance surface than the same tool implemented only in Germany. Do not leave this field empty until it is "full legal review". It's much better to have a provisional classification that is later re-classified than to leave a blank cell that never gets prioritized.
Ownership and Accountability
Each system should have a business owner (who answers for the use of the system) and a technical owner (who answers for how the system is used). “Customer Success” is not responsible for anything; the VP that signed off on the tool is! If there is a vendor contact and contract reference, include it, it's one of the few things that will naturally reveal configuration changes, due to the timing of a cycle.
Evidence Links
This is where the inventory transforms from a list to an audit-ready field! Attach a model card or vendor documentation, latest test results (bias audits, security testing, performance benchmarks), monitoring dashboards, and any existing DPIA and risk assessment. A model card is the standard one-page summary of the intended use, training data, and known limitations of a model (created as a practice in ML documentation and now close to a regulatory expectation) and is often the fastest way to answer the question "what does this system actually do and how was it validated" without involving the engineering team. It is OK to leave the row blank if there is no evidence to connect; this is a discovery in and of itself, rather than a completion failure.
Where to Find AI Systems You Don't Already Know About
According to UpGuard's State of Shadow AI report, over 80 percent of workers are using AI tools that their IT department did not approve. None of this is captured by any of the discovery methods alone, so these discovery methods are run concurrently and not sequentially.

Procurement and SSO Logs
Cross-reference your software spend from over the past 12 months with your SSO provider's application catalog for any apps labeled AI, ML, analytics, or automation. The identity platform(s) that an employee uses to sign in to any applications on their corporate device (such as Okta or Azure AD) record all the applications the employee logs into, including AI tools that were not purchased in the company's name, or that the employee signed up for via a free tier or is reimbursed for through company expenses. This one cross-reference typically yields more shadow tools than all others, as it picks up adoption that occurred independent of procurement.
Vendor Intake Questionnaires for New Tools
All new vendor relationships and renewals should include a vendor intake document that contains a specific question: Does this product use AI or machine learning as part of their service, and if so, how? This is important because there are a lot of vendors that are constantly enhancing their existing products with AI capabilities, and the enhancements do not necessarily result in a new procurement assessment. Whether or not you were told, you are now a deployer under the EU AI Act, as a payroll vendor slipped its AI-powered anomaly detector quietly into an expense-management module you've used for years.
Surveying Engineering Teams for Internally Built AI
A structured survey of engineering and data science leads, which asks each person to list all the models, scripts or automated decision systems that they have in production or staging, will reveal systems that were never recorded in a procurement log or SSO log. Bikes that pass through here, typically those that started as a side project: a scoring script that was an idea from a notebook that rarely became software, let alone AI, since it wasn't formally released. Present the survey as a governance support activity and not an audit, as teams share more when they aren't afraid of being faulted for something they did with good intentions.
Shadow AI and Shadow Agent Discovery
Shadow AI discovery and shadow agent discovery are similar, but not the same. Traditional shadow AI discovery, as it's done by CASB or SaaS-management tools, is an approach that seeks out unsanctioned web applications and browser extensions. Shadow agents are another kettle of fish: They are autonomous processes that are making API calls directly from the developer's laptop or from a CI pipeline; thus, network-layer monitoring catches only a small portion of them, because the traffic never transits through a proxy that is monitoring for them. According to IBM's 2025 Cost of a Data Breach Report, shadow AI played a role in an estimated one of every five breaches last year, and contributed an average of $670,000 to data breach costs.
Classifying What You Find
After a system is logged, it must be given a risk tier, a regulatory role, and be mapped to applicable frameworks within your organization. This is a continuous function of triage and not a sorting exercise.
Risk Tiering
The EU AI Act categorizes systems into four classes: unacceptable (those are prohibited), high-risk (Annex III, such as employment, credit, biometric identification, access to education, etc.), limited-risk (those must be made transparent), and minimal-risk. NIST does not require specific tiers, but you are expected to classify systems according to their potential impact on people and operations as part of the MAP. Most organizations stack a practical scoring model on top of both, as follows: Likelihood of harm, severity if it goes wrong, and level of autonomy the system has, all weighed into a tier which dictates the level of review depth.
Assigning Regulatory Role (Provider, Deployer, Importer)
Role assignment is not a checkbox that is ticked once on intake. The EU AI Act clearly outlines these roles, and the difference means different obligations come into effect. A provider develops a system or has developed a system placed on the market under its name, according to Article 3(3). A deployer (under Article 3(4)) is a person who uses a system created by another person, in his/her own name and authority, and in a professional setting. An importer is, as per Article 3(6), any entity from the EU who places a non-EU vendor's system on the EU market under its name. A distributor, as per Article 3(7), is not the provider nor the importer but makes a system, already placed, available at a lower stage in the chain. Beware of Article 25 "the accidental provider" rule: if you use a system under your own name, make material changes to it, or transform it into a high-risk use that it was not designed for, then you assume provider responsibilities, whether you intended it or not.
Mapping to Applicable Frameworks (EU AI Act, NIST AI RMF, ISO 42001)
Even with identical systems, the same inventory entry must be able to satisfy the different frameworks in different ways, and an inventory of AI systems for EU AI Act compliance will be different from an inventory of AI systems for NIST AI RMF compliance. In the case of the EU AI Act, it refers to the Annex III category, the conformity assessment route, and the documentation status under Annex IV, which includes nine mandatory sections, including the system description, data governance, risk management, and a post-market monitoring plan. For NIST AI RMF, it means sufficient context, purpose, data lineage, and stakeholders to facilitate the benchmarking activities of the Measure function. For ISO 42001 certification, it means that you have identified the scope of your AI management system and that each of your systems in the scope has been systematically identified and classified. One inventory, several lenses. How inventories rot within a year is by building three separate spreadsheets of three frameworks.
Structuring the Inventory
The correct inventory structure is not so important as the fact that it is a single source of truth, but it does vary in scale.
Spreadsheet/Simple Database (Early-Stage Approach)
For organizations with 50 systems or fewer and a single compliance owner, a shared spreadsheet can be a viable option regardless of whether it's an AI spreadsheet that was created in Excel or a simple Airtable spreadsheet base. Easy to rise to a standing position without having to get purchase approval for that position. The problems that can arise are well known: version conflict if two users edit it simultaneously, no tracking of the changes made, and no warnings if the review date is missed.
Dedicated GRC or AI Governance Platform (At Scale)
Moving away from a spreadsheet is typically not triggered by a set number of systems. It's when more than one business unit requires write access, when auditors begin to request a change history that you can't see in a spreadsheet's revision history, or when the review timeline requires automated reminders instead of a calendar invite you don't get to. A dedicated platform provides workflow automation for intake, built-in audit trails, and integrations with existing GRC platform tooling to avoid the AI inventory being an isolated entity from the rest of your risk register.
Recommended Tabs: Master Inventory, Discovery Log, Change Log
Regardless of the structure you select, three views perform a significant amount of the work. The Master Inventory is one row for each system with all the above system fields. The Discovery Log is used to document how and when each system is discovered, how it was discovered, and by whom, since an auditor wants to see a systematic discovery process, not a list of discovered systems. The Change Log stores all edits to existing content entries, including Risk Reclassification and Ownership Change, and includes a reason for the edit and a timestamp. With the absence of the change log, you cannot prove that the system's risk status was re-evaluated following a material modification as required by several frameworks.
Inventorying Agentic AI and MCP-Connected Systems
Not much has been written about this section, and it is now the most significant gap in AI governance initiatives. Agents don't create only outputs. They interact with actions, infrastructure, and tools that a model-only inventory could never be built to describe.
Why Agents and Tool Connections Need Their Own Inventory Fields
A static model entry gives you the prediction or generation of a system. It gives you no idea of what an agent based on that model will be able to do when it is provided with tools. The Cloud Security Alliance's AI Safety Initiative has reported that tool poisoning attacks involve an adversarial instruction being set directly into the description or parameter schema of a tool, with the agent assuming it is a trusted operating context due to the lack of a native mechanism in the Model Context Protocol to identify the tool as potentially suspicious. An agent that can access the email, code repository, and payments API does not require a new vulnerability to wreak havoc. It only requires a single poisoned tool description at initialization, and the role's ambient access takes care of the rest. An inventory entry should include an agent that can invoke it, what it can reach, and whether it must be confirmed by humans or not, which are all different from the type of risk that the model is exposed to.
Treating MCP Servers as Inventory Items, Not Just Infrastructure
MCP servers are provisioned as infrastructure, and audited as if they don't exist. That gap has already proved to be tangible, with security researchers at OX Security finding that there was a system-wide issue with how the official versions of the MCP SDK manage to execute a local tool using the STDIO transport, which Anthropic has confirmed was deliberate anyway, and that STDIO’s job is for downstream developers to fix, not for the protocol. OX Security estimated that more than 150 million downloads were hit by the affected supply chain and that there were approximately 200,000 vulnerable instances in the clients Cursor, VS Code, and Claude Code. The Cloud Security Alliance's research note on the incident made this one clear: If an organization's inability to perform a full inventory of MCP connections, then it cannot determine its exposure when a flaw like this is exposed. Each MCP server entry must have its own entry in the list, separate from the agent that interfaces with it:
If that record is infrastructure metadata, buried in a config repository, security review occurs if someone thinks about asking for it. If you consider it an inventory item with an owner and a review date, it will appear automatically when a CVE-class MCP vulnerability is disclosed. A platform like Akto's agentic security platform discovers and catalogs MCP servers, the agents attached to them, and the tools that each server exposes, across cloud infrastructure, CI/CD pipelines, and employee endpoints - and this class of asset is added to the inventory the same way a new SaaS subscription would be added: automatically, rather than waiting for it to be added to a developer's local config.
Keeping the Inventory Alive
If the inventory is accurate on the day it's published but dated within a quarter, it's already done its job as an inventory, which is to provide you with an on-the-spot answer, not a historical one. This is not a one-off; you're not going to take the initiative of closing it down and putting it in a file and forgetting about it.
Quarterly Reviews and Event-Triggered Updates
Establish a quarterly baseline review as a minimum, rather than a goal, and add on top of that event-based triggers: A new system moves into production, a foundation-model version bump with an existing agent, a material change to the training data or intended purpose, a vendor notification of new AI functionality, a decommission, or an incident with any inventoried system. All of these should trigger an update to the inventory the same week that the update is released, not the next time the inventory is reviewed. Event triggers catch the changes, or the quarterly cadence catches the drift, that create new exposure.
Why a Static Inventory Fails Within Months
According to McKinsey's 2024 State of AI survey, 72 percent of businesses have already integrated AI into at least one aspect of their operations, and adoption is surging, with most currently taking place at the departmental level, rather than centrally. A new inventory created under these conditions is a picture of a business that no longer exists by the time the next business’s procurement cycle is closed. Those that maintain a comprehensive inventory view it as a living entity that integrates with procurement, vendor management, and change advisory procedures, rather than a document that is updated when an audit comes up.
A Practical First-30-Days Plan
Creating an inventory from scratch is like several months of work until you divide it up into weeks and give each one some specific tasks to accomplish.

Week 1: Pull Procurement and SSO Data
Export the entire list of your SSO provider's applications and mark as AI-related any applications with a name or type that appear to be related to AI. Extract procurement spend data for the last 12 months, across all AI and ML, analytics, and automation line items. Overlap the lists – anything that shows up on both lists is good; anything that shows up in SSO logs but not in procurement records is your first list of shadow AI.
Week 2: Survey Engineering and Business Teams
Email a brief survey to each engineering lead and business unit head to identify any AI, machine learning, or automated decision-making tool, script, or model that their team uses or has created. Invite each of two or three teams most likely to have been influenced by something informally (customer support, sales ops, data science) to follow up with a 15-minute discussion.
Week 3-4: Classify, Assign Owners, Flag Top Risks
Execute all of the systems that have been processed through the risk tiering process, assign a business and technical owner to each system, and escalate to Annex III category and/or systems in which regulated personal data is processed for expedient legal review. At the end of week four, you will not have an ideal stock. You'll end up with a spreadsheet that's filled with data and marked with obvious gaps, as opposed to a blank one.
How Akto Helps Build and Maintain an AI System Inventory
Each problem raised to date has a discovery and monitoring answer, not a paperwork answer-that is, a paper answer. That's what Akto is meant to do.
Automated Discovery Across LLMs, Agents, and MCP Servers
Akto's agentic security platform identifies and classifies LLMs, AI agents, MCP servers, and the tools and resources each of them interacts with, using over 80 connectors to cloud infrastructure, CI/CD pipelines, and employee endpoints. What that coverage means is that any agent that an engineer wires in over a weekend, or an MCP server that's spun up in a staging environment and never used, will appear in the discovery's inventory, rather than being hidden until it's discovered by an audit.
Continuous Risk Scoring and Classification
Since they rely on live prompts, tools, and context rather than a fixed model artifact to determine their behavior, a risk tier assigned at intake and never revisited is a burden every time the underlying system changes. Akto tests for prompt injection, tool misuse, and policy bypass mapped to the OWASP LLM Top 10 and MITRE ATLAS, and it's an automated process that continues to run against discovered agents and MCP servers to provide an ongoing agentic security posture score, not a one-off assessment. This makes the classification step described above somewhat useful because it can run it on the behavior of a system today, rather than the behavior at the time of a system's initial log.
Audit-Ready Reporting
A posture score that is not backed up by a trail doesn't hold up when an auditor or customer's security team asks for evidence. While Akto's discovery and testing capabilities are enhanced by policy enforcement and runtime protection logs, it can also identify PII, credentials, and other sensitive data traversing discovered APIs and agents - the kind of documentation that the Evidence Links field in your inventory ought to be documenting.
Final Thoughts on AI system inventory
There is a simple, true answer to how to build an AI system inventory for compliance: it's not a document; it's a discovery habit: the real risk is in the shadow tools, agents, and MCP servers that people have not formally approved. Having created a spreadsheet once and looking at it during an audit will never be ahead of that reality. Agentic security gaps, no-owner agents and MCP servers, risks that do not accurately reflect current activity, and evidence collected on the fly under a deadline- those are the things Akto's agentic security platform was designed to address. The discovery process is automated and is available for over 80 connectors, so the inventory stays up to date without a quarterly scramble, and the risk scores are kept in check during red teaming periodically. When there are still more holes than answers, schedule an Agentic security demo and experience what Akto brings to the table for you.
Frequently Asked Questions
What is an AI system inventory, and why is it required for compliance?
An AI system inventory is a unified list of all the AI systems, models, and agents an organization develops, purchases, or utilizes, including who owns them, their function, and a risk assessment of their respective type. It's required because all major frameworks, including the EU AI Act, NIST AI RMF, and ISO 42001, presume that you can identify all the systems in scope without first doing any classification and documentation.
What fields should an AI system inventory include?
At minimum: core identification (name, vendor, version, deployment date), a plain language description of function, the tools and systems it's connected to, regulatory classification (risk tier, role, jurisdiction), a named business and technical owner, and links to supporting evidence such as model cards and test results. For a detailed list of what to capture, see the above section.
How do you find AI systems you don't already know about?
Cross-reference SSO application logs against procurement records, send vendors a direct questionnaire asking whether AI is used in service delivery, survey engineering and business teams directly, and run continuous technical discovery for AI API traffic and MCP connections rather than a single sweep.
How does shadow AI discovery fit into building an AI inventory?
Shadow AI discovery is the systems that have been used by individual employees or teams without a procurement process. It should be done on an ongoing basis, not just a single event, as new shadow tools will be introduced faster than most reviews can keep up.
What's the difference between an AI system inventory and an AI risk assessment?
In the inventory, what exists, who owns it, and what it's connected to is the catalog. The risk assessment is what you do with each entry after you log it: evaluating the likelihood of harm and the severity of the harm, and determining a tier. A risk assessment cannot be carried out on a system that doesn't exist in the inventory; that's why the inventory is first. Content detail is in AI Risk Assessment.
Should minimal-risk AI systems be included in the inventory?
Yes. A system that is deemed minimal-risk today can be high-risk once it has undergone a scope change or a new use case, and only be discovered by having the scope change or new use case logged from the beginning. Entries that are not considered to pose minimal risk may be documented less extensively, but no entries at all are a weak link in the inventory for the type of reclassification the regulators are looking for.
How do you classify AI systems by role (provider, deployer, importer) under the EU AI Act?
What did your organization do with the system? You are a provider under Article 3(3) if you constructed it or had it constructed and sold it under your name. If you are utilizing a system developed by another person under your own direction, you are a deployer under Article 3(4). Article 3(6): If you are an EU entity who places a non-EU vendor's system on the EU market under that vendor's name, you are an importer. If you don't mean to assume provider obligations, but you end up assuming them through rebranding, material changes, or repurposing an Article 25 system.
Should an AI inventory be a spreadsheet or a dedicated platform?
A spreadsheet is a good starting point when there are approximately 50 systems and one owner. When multiple business units require write access to a GRC or AI governance platform, auditors begin asking for a change history that a spreadsheet can't provide, or when the review cadence needs to be automated instead of by a calendar reminder, it's time to move to a dedicated GRC or AI governance platform.
How often should an AI system inventory be reviewed or updated?
Quarterly, at the very least, and event-based updates on top of that (for anything that has occurred since the last review, including a new system going into production, a material change, an AI feature added by a vendor, a decommissioning, or an incident). Use the quarterly review as a minimum, not a sole indicator.
How do you inventory AI agents and MCP-connected tools, not just models?
Provide an Agent and MCP Server inventory that isn't tied to the underlying model: list of tools that an agent is able to call, list of tools that an agent can reach, human confirmation needed for actions, and for MCP servers, which agents connect to them and what credential scope they have. Those questions are the ones that are most important when a protocol-level vulnerability is disclosed, and a model-only inventory entry will not be able to answer them.
What's the difference between building an inventory for the EU AI Act vs. NIST AI RMF?
The EU AI Act must have an Annex III category, a conformity assessment route, and an Annex IV documentation status that leads to registration and CE marking. NIST AI RMF's inventory is the same one used for feeding its Map function, which provides context and stakeholders for the Measure and Manage functions that follow. Both can be met in a single inventory if the regulatory fields are wide enough in their design.
How do embedded or third-party AI features get missed in an inventory?
Often, AI features are introduced as an add-on to existing products and systems without requiring a new procurement process because it is not a new product or system, but an update. For example, an existing payroll platform that integrates an AI-powered anomaly detection module. This can only be caught through a standing vendor questionnaire at the renewal, not necessarily at signing.
What evidence should be linked to each AI system in the inventory?
The latest security and bias testing results, monitoring or posture dashboards, any DPIA or risk assessment on file, or model cards. Leave empty systems as empty, but if there is no evidence to link, make a finding that a system is empty.
How long does it typically take to build an initial AI system inventory?
It is possible to implement a first pass of a classification, with teams surveyed and populated but not fully classified and owned, with an implementation plan in place in about 30 days, following the above process of pulling in procurement and SSO data in week one; teams surveyed in week two; and teams classified and assigned owners in weeks three and four. It's usually several months longer for full maturity, where discovery and quarterly review are automatic.
How can platforms like Akto automate AI system inventory and discovery?
With over 80 connectors, Akto automates the discovery and cataloging of LLM availability and their AI agents and MCP servers across cloud infrastructure, CI/CD pipelines, and employee endpoints, and then conducts continuous automated red teaming to ensure that risk classifications remain current, not static. This posture and runtime protection data serves as evidence for auditors to investigate.
Experience enterprise-grade Agentic Security solution

