The Shadow AI Tool Register
Answer “Which AI Tools Hold Our Data?” From One Register
Find every tool, name its owner and data, then record keep, restrict, replace, retire, or investigate.
Unlock the full document
This document is free. Leave a work email so we can send corrections and updated versions, and the full sheet unlocks below.
The client asked which AI tools hold its data
A client, auditor, carrier, board, or regulator asks whether your organization uses AI. The answer is due, and “I would have to check” reveals that no register exists.
This document is for the Solo IT Director, CTO or VP Engineering, practice administrator, firm administrator, or managing partner who owns that answer. Find every tool in use, its owner, the data entered, the vendor’s data treatment, and the written decision for that tool.
Build the register in a spreadsheet. Buy nothing to start.
Count the tools before you approve another platform
Inventory comes before procurement.
You cannot write an AI policy for tools you have not found. You cannot pick an approved platform before you know what people are actually doing with the unapproved ones. And you cannot buy a governance product to solve a problem you have not measured, though several will be sold to you on exactly that basis.
If your first inventory finds 14 unapproved tools beside one licensed platform, the policy did not produce a usable decision record. Count the tools, name their owners and data, and decide each row.
Find first. Decide second. Buy last, if at all.
Use your own tool count in the decision record
There is a widely quoted figure about employees bringing their own AI tools to work. Here is the real thing, with its limits attached.
In a 2024 survey of about 31,000 workers across 31 markets, 75% reported using AI at work. Among those AI users, 78% reported bringing their own AI tools. The result is self-reported, the sample is a vendor survey, and it does not establish prevalence inside your organization in 2026.
Treat that as a reason to look, and nothing more. Your own count is the only number that matters, and it is the number you are about to produce.
Part 1: Run six searches for tools already in use
Six methods. Run all six. Each one finds a category the others miss, which is the reason to run all six rather than picking the easiest.
| Method | What it finds | What it misses | Effort |
|---|---|---|---|
| Expense reports and corporate cards | Paid personal subscriptions, usually $20 to $60 per month | Free tiers, personal cards, anything expensed as “software” without detail | 2 hours |
| Accounts payable and invoice review | Departmental purchases, annual contracts, tools bought by marketing or ops without IT | Anything under an approval threshold, anything on a credit card | 2 hours |
| Identity provider and SSO logs | Anything staff signed into with a work account, including OAuth grants to your tenant | Tools accessed with personal email addresses | 3 hours |
| Browser extension and app inventory | Extensions, plugins, desktop clients, IDE add-ins | Anything used entirely in a browser tab | 3 hours |
| SaaS discovery tooling | Sanctioned and unsanctioned SaaS across the fleet, if you already own such tooling | Personal devices, personal accounts, one-off web use | Varies |
| Asking people directly | Almost everything else, and the reasons behind it | Nothing, if you run it correctly | 1 week elapsed |
Use a two-week amnesty to collect the missing names
Asking people finds more than every technical method combined. It also fails completely if you run it as an enforcement action.
Run it as an amnesty. Say so in writing, in these terms or your own:
We are building a list of the AI tools people use for work. This is an inventory, not an investigation. Nobody will be disciplined for anything they report in the next two weeks. We need to know what is in use so we can tell clients and auditors the truth, and so we can make sure the tools people rely on are ones we can keep.
Then hold to it. If one person gets reprimanded during the amnesty window, the inventory is over and every future inventory is over too. Word travels faster than your survey.
Ask four questions and nothing more:
- What AI tools do you use for work, including free ones and ones on your personal account.
- What do you use them for.
- What information do you put into them.
- What would break for you if it went away tomorrow.
The fourth question is the one people answer honestly, and it tells you which tools have become load-bearing. A tool that would break a workflow if removed is not a shadow tool anymore. It is unmanaged infrastructure.
Pull OAuth grants that can read mail, files, or calendars
If you run a Microsoft or Google tenant, pull the list of third-party applications that hold OAuth consent against it. This is a directory query, it takes twenty minutes, and it regularly surfaces AI tools that were granted read access to mail, files, or calendars by a single user click. Those grants persist after the person stops using the tool. They persist after the person leaves.
Part 2: Give every tool 11 decision fields
Eleven fields. Every one of them exists because a specific question gets asked later and the answer has to come from somewhere.
| Field | What goes in it | Why it exists |
|---|---|---|
| Tool | The product name and the tier or plan in use | Free tier and paid tier of the same product often have different data terms |
| Owner | A named person, never a department | Somebody has to answer questions about this tool. If no name fits, that is your first finding |
| Business use | The actual task, in one sentence | Distinguishes a tool that drafts internal email from one that screens job applicants |
| Data entered | The classes of information that go in: client names, health information, financial records, source code, contract text, nothing sensitive | This is the field that determines everything downstream |
| Model or provider | Who runs the model behind the tool, if the tool discloses it | Many AI products are interfaces over someone else’s model. Your data may reach a party you have never heard of |
| Retention | How long inputs and outputs are kept, and where that is stated | If the answer is “unknown,” write unknown. Unknown is a finding |
| Access | Who inside your organization can use it, and who at the vendor can see the data | Both halves matter |
| Contract | None, click-through terms, signed agreement, or enterprise agreement with a DPA | A click-through consumer agreement is a real contract, and it is usually a bad one |
| Risk tier | 1, 2, or 3, defined below | Determines review depth and review frequency |
| Approval | Approved, restricted, replace, retire, needs evidence, with a date and a name | The decision record |
| Review date | When this row gets looked at again | AI products change terms and models frequently. A register with no review dates goes stale in a quarter |
Put each tool in one of three review tiers
Keep this to three tiers. More tiers produce debates instead of decisions.
| Tier | Definition | Review frequency |
|---|---|---|
| Tier 1 | Regulated or client-confidential data goes in, or output drives a consequential decision about a person, or the tool touches source code or production systems | Quarterly, plus review on any vendor term change |
| Tier 2 | Internal non-public business information goes in, and output is reviewed by a person before use | Twice yearly |
| Tier 3 | Public information only, output is drafting or ideation, no decision authority | Annually |
Part 3: Screen the data, contract, owner, and client exposure
Six questions per tool. They are yes or no. Answer them from evidence, not from what the marketing page says.
| # | Screen question | Why it matters | If yes |
|---|---|---|---|
| 1 | Consequential decision use. Does output influence a decision about a specific person: hiring, promotion, termination, credit, insurance, housing, education, or access to a service or treatment | This is the category regulators are moving on first, and it is where a wrong answer has the widest effect | Escalate to counsel before proceeding. Document the human review step in writing |
| 2 | Human review. Does a qualified person review the output before it is acted on, and can that person tell when the output is wrong | An output nobody checks is a decision the tool is making | If no, this is your highest-priority finding regardless of tier |
| 3 | Sensitive data. Does regulated data enter the tool: health information, financial account data, Social Security numbers, biometric data, or anything covered by a confidentiality obligation | Determines whether contract terms and a data agreement are required rather than optional | Requires a signed agreement with data terms. Click-through terms do not carry you here |
| 4 | Customer data. Does information belonging to your clients or customers enter the tool | Your client contracts may already prohibit this or require notice and consent | Check your client agreements before anything else. Some prohibit subprocessors outright |
| 5 | Source code use. Does proprietary code enter the tool, and does the tool suggest code that enters your product | Two separate exposures: your code leaving, and code of unknown provenance arriving | Requires a licensing review of suggested output and a code review control |
| 6 | Training use. Does the vendor use your inputs or outputs to train or improve models, and can that be turned off | Once content is in a training corpus, no deletion request retrieves it | Turn it off in settings, get it in writing, and record where the setting lives |
Screens 1 and 2 are separate questions and must stay separate. A tool can be used in a consequential decision with strong human review. A tool can be used in something harmless with no review at all. Collapsing the two produces the wrong answer in both directions.
On the regulatory picture, stated accurately
The federal government has not imposed an AI certification requirement. NIST published a generative AI profile that gives voluntary actions for identifying and managing generative AI risk. It is guidance. It is not a certification and it carries no deadline. Anyone selling you “NIST AI certification” is selling something that does not exist.
State law is where the dates live. Colorado’s amended AI law carries an enacted effective date of January 1, 2027 for covered automated systems used in consequential decisions. Whether your systems fall in scope is a legal determination that your counsel makes, not one this document can make for you.
Separately, state privacy laws continue to take effect on staggered dates: Louisiana and Oklahoma on January 1, 2027, Alabama on May 1, 2027, Vermont on January 1, 2028. Applicability depends on each law’s own thresholds and exemptions, which vary by state, entity type, and data type.
What all of this means practically: screen 1 exists because the obligations attached to consequential-decision systems are the ones with real dates on them. Get counsel involved for any tool where screen 1 is yes. For everything else, the register itself is the work.
Part 4: Record keep, restrict, replace, retire, or investigate
Every row ends in one of five decisions. No row stays blank.
| Disposition | Means | Requires before you record it |
|---|---|---|
| Approve | Continue use as-is | Named owner, data classes recorded, training use confirmed off where applicable, review date set |
| Restrict | Continue use with named limits | The limits written down in a sentence a user can follow, and the person who enforces them |
| Replace | The use case is legitimate, this tool is the wrong vehicle | The named replacement, the migration owner, and a date. A replacement with no date is a retire in slow motion |
| Retire | Stop use | An owner, a date, and a check that nothing depends on it. Run question four from the amnesty survey again before you pull it |
| Needs evidence | The decision cannot be made yet | The specific missing fact, who is getting it, and by when. This is a legitimate state and a temporary one |
Needs evidence is the honest answer for most rows on day one. That is fine. What is not fine is a register still full of needs evidence ninety days later, because at that point it has become a record of things you decided not to decide.
A useful discipline: cap needs evidence at thirty days per row. When the clock runs out, the row becomes restrict by default until the evidence arrives. Restriction is reversible. Silence is not a decision.
Part 5: Review one illustrative AI tool register
The tools below are described generically on purpose. Naming platforms would make this document a recommendation, and it is not one. Substitute the real product names when you fill your own register.
Record the tool, owner, users, cost, and approval status
| # | Tool | Owner | Business use | Data entered |
|---|---|---|---|---|
| 1 | General-purpose assistant, consumer free tier, personal accounts | Unassigned at discovery, now: Dir. of Operations | Drafting client emails and summarizing documents | Client names, matter details, occasional contract text |
| 2 | Meeting transcription and summary service | Sales Manager | Recording and summarizing client calls | Full call audio, client names, pricing discussions |
| 3 | Code completion plugin, individual licenses | VP Engineering | In-editor code suggestion | Proprietary source code, repository context |
| 4 | Marketing copy generator | Marketing Manager | Blog drafts and social posts | Public marketing material only |
| 5 | Resume screening and ranking feature inside the applicant tracking system | HR Manager | Ranking applicants for review | Applicant names, employment history, education |
| 6 | Document analysis assistant, browser-based | Paralegal (self-adopted) | Summarizing discovery documents | Client-confidential litigation material |
| 7 | Customer support reply drafter, built into the helpdesk platform | Support Lead | Suggesting first-draft replies | Customer names, ticket history, account identifiers |
Record the data, training use, retention, access, and contract terms
| # | Model or provider | Retention | Access | Contract | Tier | Screens triggered | Disposition |
|---|---|---|---|---|---|---|---|
| 1 | Undisclosed by tool | Unknown | 11 staff, personal accounts | Click-through consumer terms | 1 | 3, 4, 6 | Replace by Nov 30. Named business-tier replacement with data terms. Owner: Dir. of Operations |
| 2 | Third-party model, named in docs | 90 days, stated in settings | Sales team, 6 users | Paid plan, click-through | 1 | 3, 4, 6 | Restrict. No recording of calls involving health or financial account data. Consent notice added to invite text |
| 3 | Vendor-run model | Snippets retained per enterprise setting | Engineering, 14 users | Business agreement in place | 1 | 5, 6 | Approve. Training on private code confirmed off in tenant settings, screenshot filed. Code review control already required |
| 4 | Undisclosed | Unknown | Marketing, 3 users | Click-through | 3 | None | Approve. Public data only. Annual review |
| 5 | Feature of existing HR platform | Per master agreement | HR, 2 users | Signed MSA, no AI-specific terms | 1 | 1, 2 | Needs evidence to Oct 15. Counsel review of consequential-decision scope. Human review step being documented. Vendor asked for model and validation documentation in writing |
| 6 | Undisclosed | Unknown | 1 user | Click-through consumer terms | 1 | 3, 4, 6 | Retire immediately. Client engagement letters prohibit disclosure to third parties without consent. Replacement request opened |
| 7 | Undisclosed by helpdesk vendor | Unknown | Support, 9 users | Existing platform contract, AI feature added by vendor in a release | 2 | 4, 6 | Needs evidence to Oct 15. Vendor asked whether the AI feature introduces a subprocessor and whether ticket content trains models |
Use the open fields to assign the next decision
Row 7 is the one worth studying. Nobody adopted that tool. The vendor shipped an AI feature into a platform the organization already licensed, and the data flow changed without any purchase, any decision, or any notice that reached IT. Your purchase log will not show this route. Add release-note review and a written vendor question to the register process.
Add a standing item: when any existing vendor announces an AI feature, the register gets a row.
Row 5 is the second one worth studying. It triggered screen 1, so it stopped being an IT decision.
Part 6: Issue one page while open tool decisions are resolved
Adapt this. Do not adopt it as-is, because your regulatory obligations, client contracts, and risk appetite are yours. Have counsel review before publication. The value here is the shape and the length: one page that a person will actually read, written in sentences a person can follow.
[Organization] AI Tool Use Standard Effective [date]. Owner: [name, title]. Next review: [date].
Why this exists. AI tools can help us work faster. They can also send information outside the organization in ways that violate our client commitments and our legal obligations. This standard says which tools you may use and what you may put into them.
Scope. This applies to every AI tool used for [Organization] work, including free tools, tools on your personal account, tools on your personal device, and AI features added to software we already use.
The approved list. The current list of approved tools is at [location]. If a tool is not on that list, it is not approved for work use yet. Ask [name] and you will get an answer within [N] business days. Asking is never held against you.
What you must never enter into any AI tool without written approval:
- [Client or patient names and any information that identifies them]
- [Health information]
- [Social Security numbers, financial account numbers, payment card data]
- [Passwords, keys, tokens, or credentials of any kind]
- [Proprietary source code]
- [Contract text, litigation material, or anything covered by privilege or confidentiality]
- [Anything you would not send to a stranger by email]
Human review. You are responsible for anything you produce with an AI tool, the same as if you wrote it yourself. Check the output. AI tools produce confident text that is wrong, including invented citations, invented figures, and invented sources. If you cannot verify it, do not send it.
Decisions about people. Do not use an AI tool to make or materially influence a decision about hiring, promotion, discipline, termination, credit, insurance, treatment, or access to a service, unless [name] has approved that specific use in writing.
AI features in existing software. If a vendor turns on a new AI feature in software you use, tell [name]. You do not need to evaluate it. Just report it.
If something goes wrong. If you entered information you should not have, tell [name] the same day. There is no penalty for reporting it promptly. There is a real problem if we find out later from someone else.
Questions. [Name], [contact].
Two notes on the wording above. The reporting clause is deliberately protective, because an organization that punishes self-reporting learns about its exposures from clients instead of from staff. And the “asking is never held against you” line does real work, since the fastest route back to shadow AI is a request process that feels like a trap.
Part 7: Price a governance product only after you know the open rows
Before anyone quotes you a governance platform, sort your findings into three groups.
Group A: costs nothing but attention. Turning off training use in settings you already have. Assigning owners. Revoking stale OAuth grants. Writing the one-page standard. Retiring tools nobody depends on. Recording data classes. Most organizations find that half or more of their register work lands here.
Group B: costs configuration or a plan change. Moving people from consumer tiers to business tiers of tools they already use, which is often where the data terms actually change. Enabling tenant-level controls in the identity platform you already pay for. Documenting a human review step.
Group C: costs money. Discovery tooling. A governance platform. A new approved platform to replace several unapproved ones.
Do A and B before anyone prices C. And apply one test to any Group C proposal: ask which of your register rows the product would have found, and which of your dispositions it would have decided. A product that produces another inventory you still have to review by hand is duplicating work you have now already done.
If a vendor’s first response to your register is a platform proposal, ask which of your findings are in Group A, and why those were not the recommendation.
Add every new feature, vendor, and review date after the first pass
The first pass is a project. What follows is a habit, and the habit is short.
| Trigger | Action |
|---|---|
| New tool request | New row before approval, never after |
| Existing vendor announces an AI feature | New row, screens 4 and 6 at minimum |
| Quarterly, Tier 1 rows | Re-check retention and training settings. They change without notice |
| Any vendor terms-of-service update | Re-check contract and retention fields for that row |
| Staff departure | Check owned rows and reassign. An unowned Tier 1 row is an open finding |
| Client questionnaire or audit request arrives | The register is the answer. This is the payoff |
The last row is why the register is worth building even if nothing on it is risky. When a client asks whether you use AI on their matter, the difference between a defensible answer and a guess is this spreadsheet.
Keep the register ready for the next client request
This document is general guidance on inventorying and governing AI tool use. It is not legal advice. Whether a specific tool, use case, or data flow triggers an obligation under state AI law, state privacy law, HIPAA, GLBA, professional responsibility rules, or your own client agreements is a legal determination. Have your counsel make it.
Every regulatory and statistical claim above is sourced inline to a named authority through SBK’s source register. Where a requirement varies by state, entity type, or data type, this document says so rather than giving you one number. Where the honest answer is that the evidence does not support a claim, it says that too.
SBK Consulting is a family-run, vendor-neutral IT advisory firm serving the New York, Connecticut, and New Jersey metro area. Founded 2010. Zero vendor partnerships, zero reselling, zero commissions or referral fees. 125+ years combined experience, 100% US-based. We do register reviews if you want a second set of eyes on your Tier 1 rows and your needs evidence list. If you never call us and this document gets your AI use written down, it did its job.
(718) 407-4169