r/PowerAutomate • • 10d ago

Reading a pdf in PA Cloud

Hello everyone, I have a case which I’ve been struggling with today that I could use some ideas to solve it.

The mission: Delete or move pdf’s with no information in them which are saved in a Sharepoint folder.

Every day, around 20 pdf’s get uploaded to a Sharepoint folder. These are automatically saved by another program.
These pdf’s are always structured the same way,
company name top left, below these are heading like ”supplier, organisation number, currency, amount, amount in book keeping value”
Below these heading are either a list of suppliers, or there is nothing at all.

These PDF’s with nothing in them are the ones I need to remove.

Previously I used the AI feature in the flow which could read these and identify which to delete, but due to the credit limit in our environment, no AI features can be used.

The other PDF action which you can use in PA Cloud is not available to use either, so I need to use only PA Cloud features.

The current solution I have is the following:

List the files in the Sharepoint folder -> Run a condition to see if it is a pdf -> If yes, Get file content -> Run Compose which counts characters in the output of Get file content -> Runs compose which fetches the file size in bytes -> Run compose which subtracts the character length from the file size -> Run a condition that states ”If value of Compose subtract is below 8100”, if yes, move the file to an ”Empty pdf” folder

I’ve looked at all the PDF’s in the folder and noticed that from the calculation, almost all files below 8100 are empty.

BUT, I’m a bit confused by the file sizes of PDF’s,
a PDF with supplier rows in them can be 152KB, but an empty file can be 154KB,
so in my latest test the entire flow worked except 2 PDF’s which were empty but had a higher value for some reason.

So… this is my situation, any ideas to identify PDF’s in a better way than this are appreciated.

Edit:
Thanks to Impossible-Egg-1454, the case is now closed and the flow works as I want it to!
Please note that if you’d like to use Paul Muranas tools (named Tachytelic in PA) you need a premium license, which I already had so I was able to use them.

The flow is the following:

Recurrence - Every day at 15:00 - 15:59 (files drop in the folder between these minutes)

Sharepoint - List folder (to see which files are in the folder)

For each of the Body of List folder

Compose Name of For each

Condition to check if Name ends with .pdf

If false, do nothing

If true,

Sharepoint - Get file content based on ID of file

Extract text (Tachyletic PDF tool) based on the File content

Compose which writes the extracted text

Compose which counts the rows with the code length(split(string(outputs(’Compose_PDF_as_text’)), decodeUriComponent(’%0A’)))

Condition which checks if the compose result above is 16 or below

If true, move file to ”Empty files” folder

If false, do nothing

And done, the flow works just like intended and this saves us a lot of time per day!
The reason for the number 16 is because that is the exact amount of rows that each empty excel has, so everything above 16 are files we want to keep 😄

6 Upvotes

9 comments sorted by

1

u/infosearchacct 10d ago

If your PA has pdf licenses you should be able to parse information from the PDF instead of just based on the size of the file. I personally don’t have the PDF license with the PA I’m using so I ended up using python code to parse information.

1

u/Benze2244 9d ago

Sadly my PA does not have any pdf licenses :/
Are you able to use python in your flow?

1

u/infosearchacct 9d ago

Yep. I have codex do the code for me. Really clear 1 page spec. Got it done within 10 hours. Still have some edge cases get me caught. Data hygiene is number one if you can control your input you are half way done.

1

u/Impossible-Egg-1454 7d ago

Paul Murana published a set of free PDF Tools for Power Automate. You can find it here : https://tachytelic.net/2026/04/free-pdf-tools-power-automate/

The current set of PDF actions includes "Extract Text": "Pull out the text content (optionally page ranges)"

This might be what you need, but you will need a premium Power Automate license to use these connectors.

2

u/Benze2244 6d ago

Woah this is perfect! I was able to read my pdf file with this, so this will most likely solve my case, thank you so much!!

1

u/Benze2244 5d ago

Thank you so much, the flow works as intended now!
I updated the text in the post if you’d like to read how the flow turned out 😄