Most of the AI exposure in a device company right now sits outside the device. Here's where it shows up in your quality system, and what to put in place before somebody else finds it first.
Somewhere in your design history file is a requirement that an AI wrote.
Maybe a hazard description. Maybe the first draft of a test protocol. Maybe a whole paragraph of a risk analysis that an engineer pasted in at 11 PM last March, read twice, cleaned up, and moved on.
Nobody logged it but nobody hid it, either. There was just no field for it.
There are teams spending a lot of time arguing about how to validate a machine learning model inside a software as a medical device (SaMD) product. Those are good arguments, and they're necessary ones.
And while that argument is happening, the AI is touching the quality system every single day, as it's drafting procedures, summarizing complaints, writing code, suggesting hazards, in the form of a corporate card and a browser extension.
That second category is what I want to talk about. The tools that are already in the building. The question an investigator eventually asks is how the work got done and who was accountable for it, and that question likely has a process answer, not a technical one.
If you want the regulatory picture first, our piece on where the Food and Drug Administration (FDA) stands on AI in medical devices covers what's been published and what's still in draft. Start there if the guidance itself is new to you. Everything below assumes you've already made peace with the regulations and you're stuck on the operating part.
→ BONUS RESOURCE: FDA guidance on AI-enabled devices
Every one of them costs you nothing on the day it happens.
An engineer drafts faster, or a quality manager gets a complaint summary in 30 seconds instead of 30 minutes. The tool is good, the people are careful, and the output is usually better than the first draft a tired human would have produced at the same hour.
So there's no failure to investigate, no nonconformance, and no bad week that makes somebody ask a hard question.
The bill arrives about 18 months later, in an audit room, when an investigator picks a record and asks how it was produced. And the honest answer is that 6 different people were making 6 different judgment calls about a tool that no procedure in your quality management system (QMS) has ever mentioned.
You'd never let a device into design control without an intended use statement. That sentence is what makes every other decision possible: risk classification, verification scope, acceptance criteria, labeling.
Then teams drop a large language model into the middle of their documentation workflow and write nothing.
Without an intended use for the tool, you can't tier its risk. You can't scope its validation. You can't write acceptance criteria, because you never said what "acceptable" means. ISO 13485 clause 4.1.6 asks you to validate computer software used in the quality management system, proportional to the risk. Proportional to the risk of what, exactly? You never said.
So here's what you should do. Write one sentence per tool, that sounds something like this:
This tool produces first-pass draft text for review by a qualified engineer. It does not approve documents, release records, or make accept and reject decisions.
That sentence only takes a few minutes to write, but it does an enormous amount of work. It scopes your validation. It tells the user where the line is. It gives an auditor something to audit against, which is a much better position than watching them invent their own standard on the spot.
Then keep an inventory. Tool, owner, intended use, risk tier, version, all in one table. If you can't produce that table, that should be the first thing you do.
This is the most-claimed and least-documented control in medtech right now.
Ask 4 questions about it and watch what happens:
Most teams get through question 1. Some get through question 2. Almost nobody has an answer for question 4 that isn't "well, they approved the document."
There's a second problem underneath, and it's a human factors problem. AI output reads well (most of the time). It's fluent, confident, formatted, and arrives already looking like a finished document, versus a junior engineer's rough draft that typically gets torn apart in review. The same content, delivered in clean prose with confident hedging, gets a lighter read. That's automation bias, it's well documented, and your review control is built directly on top of it.
The fix is to separate 2 kinds of use and govern them differently.
Assistive use is where the AI produces something a qualified person evaluates, edits, and owns. The person's judgment is the control. Your record needs to show the person was competent and the review was real.
Determinative use is where the AI output drives a decision with little practical human scrutiny. That needs to include triage, classification, and scoring. For example, a complaint summary that determines whether something gets escalated. Here the tool is functioning as part of the process, and it needs validation, monitoring, and a documented failure mode analysis, because the human is no longer the control.
Teams get in trouble when something migrates from the first category to the second without anybody noticing. That migration happens because the tool got good and people started trusting it, which is the most natural thing in the world and also when your governance is most likely to have gone stale.
Let's say you qualified the tool's behavior in March, and the vendor pushed a model update in June. If nobody told you it's because nobody had to.
Which means your validation record describes a thing that no longer exists.
This is a supplier control problem and a change control problem, found in clauses 7.4 and 7.3.9, and it doesn't really have anything to do with your validation methodology. You can run a beautiful validation protocol against a moving target and still end up with a record that's not worth anything 90 days later.
Four things fix this situation:
The Purolea warning letter is the case study to read alongside this one. That piece walks through the validation evidence the FDA actually expected to see. What I'd add here is that the underlying failure was procedural before it was technical. Nobody had written down what the tool was for or what would trigger a fresh look at it.
→ BONUS RESOURCE: What the Purolea warning letter means for AI in medtech
Your traceability captures outputs. It rarely captures decisions.
When a design review happens in a conference room, the rationale ends up in the minutes. When a prompt replaces that conversation, the rationale ends up in a chat thread on somebody's personal account, and 2 years later you're trying to reconstruct why a specific failure mode was scoped out of the hazard analysis.
The engineer who did it is at a different company now, and the thread is gone. The hazard analysis just says what it says.
There are two things you can do to keep this from happening.
First, the reasoning goes in the controlled record, written by the human, in their words. What was considered, what was rejected, why. The AI can draft that paragraph, but the person has to own it and sign it.
Second, don't let the transcript become the record. When a chat log is your evidence of a quality decision, it's a quality record, which pulls it under 21 CFR Part 11 for audit trail, access control, and retention. Most chat interfaces meet approximately none of that. Keep the decision in the system that was built to hold decisions, and let the chat be scratch paper.
ISO 13485 clause 6.2 asks you to determine competence, provide training, and keep records. Teams roll out an AI tool with a Slack announcement and a link.
So what if somebody pastes a complaint narrative with patient identifiers into a public model? And then somebody else uploads an unreleased design output into a tool whose terms grant a license to the input? Neither person did anything malicious. It's just that nobody had ever told them where the boundary was, because nobody had drawn one.
What this needs is boring and specific:
If your AI tools aren't in your training matrix, they aren't governed, they're just popular.
Each of these is the same failure through a different lens.
The tool entered the company as a productivity decision. Somebody expensed it, somebody liked it, and it spread. It never entered as a controlled item, so your quality system never reached for it. And as a result it has no intended use, no owner, no version, no training, no supplier file, and no change control.
Your QMS is very good at governing things it knows about, but it can't do anything about the tools that come through the side door.
Calendar dates aren't good triggers. Events are decent. Each of these 5 is a signal that your AI governance needs to change sooner than later.
When a second person starts using the tool for regulated work. One careful engineer holding good judgment in their head is a functioning control. Two people with different standards, however, is a process that hasn't been written down.
The first time AI-assisted output lands in a controlled record. That's when the tool is inside your QMS whether you've acknowledged it or not, and that's when it's time to have appropriate intended use statements, inventory entries, and a defined owner.
When the vendor changes terms, pricing tiers, or model versions. Read the notice, don't delete it. That email is a supplier change notification even though that's not what it's called.
When output starts influencing a decision with patient consequence. Complaint triage, risk scoring, adverse event classification, anything reaching a reportability call. This is the migration from assistive to determinative, and it requires a different level of evidence.
Before your next audit or inspection, run the trace yourself. Pick 2 records you know had AI involvement and follow them all the way. Ask who used what tool, at what version, under what intended use, reviewed by whom, against what criteria, with what record. Whatever breaks in your mock trace is what will break in the real one.
That last one may take a few hours, but it's the highest-value afternoon your team will spend this quarter, and the cost is only the willingness to find out.
Every AI tool touching regulated work has an owner, an intended use, a risk tier, and a version recorded in the same place as everything else. Reviews of AI-assisted output have named criteria and leave a record. Supplier files exist for the AI vendors, with change notification in writing. The training matrix includes the tools and the boundaries. And when an auditor asks how a document was produced, the answer is easily found through a search rather than an excavation.
You can build all of that on a shared drive and a spreadsheet. Companies have done this. It costs roughly 1 full-time equivalent in administrative work, and it degrades the first time the vendor changes something and the spreadsheet doesn't notice.
Greenlight Guru builds quality management software for medical device companies specifically, which means supplier records, training assignments, design controls, risk, and document control reference each other instead of sitting in 5 tools with a spreadsheet stapled between them. When a supplier record changes, the connected requalification work follows it. When somebody asks who reviewed what, against what criteria, that's easily queried.
If AI is already in your development workflow and none of it appears anywhere in your quality system, that's worth more than 30 minutes of your attention this week.
If you are building out your AI governance process, these related guides go deeper on the specific pieces:
→ FDA guidance on AI-enabled devices
→ What the Purolea warning letter means for AI in medtech
→ CSV vs. CSA: FDA software validation guidance
→ QMS software validation under ISO 13485:2016
→ 21 CFR Part 11: a guide to FDA's requirements
→ AI, automation, and risk in medtech
See how Greenlight Guru works and bring your current tool inventory. We'll show you where it fits.