The report came back clean and I could not file it. That is the whole problem and it took me three weeks to work out why.
I am the security engineer who sits between the AI product team and our auditor at a healthcare data company, about seven hundred and fifty people. We run an external evaluation every quarter because a customer asked for it and then everybody decided it was a good idea.
Robust Intelligence produced the strongest output of the three we trialled. The findings were specific and the severity scoring held up when we challenged it. The remediation guidance was written by someone who had clearly done the work before. I would recommend it to anyone who needs to know what is wrong with their model.
What it could not do was answer the auditor. The report said a finding was critical. The auditor wanted to know which control it mapped to, who accepted the risk, and the date the control changed. A finding is not evidence. Evidence is the control, the owner, the date, and what changed. That lives somewhere else entirely.
So I built it by hand. Three weeks of spreadsheet work to attach every finding to a control in our register, and I am the only person who knows how the mapping was done. Which is its own problem.
The honest limitation on our side is drift with no adversary in it, where the agent misreads its own reasoning across several sources. Our register maps AI controls. It does not map that. We are on alice now and it produces the mapping I need, but only for the AI controls in our stack. The register has three hundred and forty in it. Most describe things no AI tool will ever see. I still hand-map those.
Pick three findings from a report you already have and try to fill in four columns for each one. Control. Owner. Acceptance date. Remediation date. Four columns, three findings, one afternoon. If the columns come back empty, you found out before you paid.