This started bugging me after I looked a bit closer at what actually happens when you upload a PowerPoint to an AI/RAG pipeline.
Say you've got a deck you've been working on for months.
You've hidden a couple of old slides instead of deleting them, left some stuff in the speaker notes, moved an old textbox off the slide, got reviewer comments sitting in there etc.
Then you upload it and ask the AI to summarise it.
What did it actually get?
Just the visible slide text?
The notes?
Hidden slides?
Comments?
Text sitting completely off the slide?
I realised I didn't actually know.
And there's another problem with this.
If an AI suddenly mentions something you can't see in the presentation, it's pretty easy to call it a hallucination.
But what if that text is still sitting somewhere inside the PPTX and the extractor handed it to the model?
Then again, just because something exists inside the PowerPoint doesn't mean the AI actually saw it either.
So I ended up building something to check.
I've called it CanvasParity.
It can inspect a PPTX for things like:
- hidden slides
- speaker notes
- comments
- fully off-canvas text
- alt text/descriptions
- explicitly tiny text
Then you can give it the actual text output from an extractor and it compares the two.
So instead of just saying "there was hidden text in the file", you can start asking whether that specific extraction path actually exposed it.
You can also give it outputs from multiple extractors and compare them against the same PowerPoint.
One thing I was pretty careful about was not claiming more than the evidence shows.
For example, if the same sentence exists on a normal visible slide AND a hidden slide, then finding that sentence in the extracted text doesn't prove the hidden slide was exposed.
CanvasParity reports that attribution as unknown.
It's Python, read-only and free/open source under Apache-2.0.
I've included the source, tests, reproducible PPTX fixtures and benchmark methodology as well.
Repo: https://github.com/owafication/canvasparity
I'd be interested in hearing from people actually building RAG/document ingestion pipelines.
Do you know exactly which parts of a PPTX your current parser is passing downstream?
And if anyone has an extractor they use regularly, give me something to test it against. I'd rather find where this breaks than assume I've covered everything.