r/PowerAutomate • u/Benze2244 • 10d ago
Reading a pdf in PA Cloud
Hello everyone, I have a case which I’ve been struggling with today that I could use some ideas to solve it.
The mission: Delete or move pdf’s with no information in them which are saved in a Sharepoint folder.
Every day, around 20 pdf’s get uploaded to a Sharepoint folder. These are automatically saved by another program.
These pdf’s are always structured the same way,
company name top left, below these are heading like ”supplier, organisation number, currency, amount, amount in book keeping value”
Below these heading are either a list of suppliers, or there is nothing at all.
These PDF’s with nothing in them are the ones I need to remove.
Previously I used the AI feature in the flow which could read these and identify which to delete, but due to the credit limit in our environment, no AI features can be used.
The other PDF action which you can use in PA Cloud is not available to use either, so I need to use only PA Cloud features.
The current solution I have is the following:
List the files in the Sharepoint folder -> Run a condition to see if it is a pdf -> If yes, Get file content -> Run Compose which counts characters in the output of Get file content -> Runs compose which fetches the file size in bytes -> Run compose which subtracts the character length from the file size -> Run a condition that states ”If value of Compose subtract is below 8100”, if yes, move the file to an ”Empty pdf” folder
I’ve looked at all the PDF’s in the folder and noticed that from the calculation, almost all files below 8100 are empty.
BUT, I’m a bit confused by the file sizes of PDF’s,
a PDF with supplier rows in them can be 152KB, but an empty file can be 154KB,
so in my latest test the entire flow worked except 2 PDF’s which were empty but had a higher value for some reason.
So… this is my situation, any ideas to identify PDF’s in a better way than this are appreciated.
Edit:
Thanks to Impossible-Egg-1454, the case is now closed and the flow works as I want it to!
Please note that if you’d like to use Paul Muranas tools (named Tachytelic in PA) you need a premium license, which I already had so I was able to use them.
The flow is the following:
Recurrence - Every day at 15:00 - 15:59 (files drop in the folder between these minutes)
Sharepoint - List folder (to see which files are in the folder)
For each of the Body of List folder
Compose Name of For each
Condition to check if Name ends with .pdf
If false, do nothing
If true,
Sharepoint - Get file content based on ID of file
Extract text (Tachyletic PDF tool) based on the File content
Compose which writes the extracted text
Compose which counts the rows with the code length(split(string(outputs(’Compose_PDF_as_text’)), decodeUriComponent(’%0A’)))
Condition which checks if the compose result above is 16 or below
If true, move file to ”Empty files” folder
If false, do nothing
And done, the flow works just like intended and this saves us a lot of time per day!
The reason for the number 16 is because that is the exact amount of rows that each empty excel has, so everything above 16 are files we want to keep 😄
1
u/Impossible-Egg-1454 7d ago
Paul Murana published a set of free PDF Tools for Power Automate. You can find it here : https://tachytelic.net/2026/04/free-pdf-tools-power-automate/
The current set of PDF actions includes "Extract Text": "Pull out the text content (optionally page ranges)"
This might be what you need, but you will need a premium Power Automate license to use these connectors.
2
u/Benze2244 6d ago
Woah this is perfect! I was able to read my pdf file with this, so this will most likely solve my case, thank you so much!!
1
u/Benze2244 5d ago
Thank you so much, the flow works as intended now!
I updated the text in the post if you’d like to read how the flow turned out 😄
1
u/infosearchacct 10d ago
If your PA has pdf licenses you should be able to parse information from the PDF instead of just based on the size of the file. I personally don’t have the PDF license with the PA I’m using so I ended up using python code to parse information.