Not sure if this is a Pickaxe problem, but I noticed that a PDF with scanned text (images) looks empty to the bots. Many documents get scanned in to PDFs meaning they aren’t “text” so to speak though have clean text. Being able to read this with the LLM would be very important and useful.
Is this something Pickaxe can fix or is this just an issue with the LLMs as whole?
Example of a chat with large document with text scanned as image: 87100041-f620-4b2d-b23c-50591a3777c7
Hi @hurmuli, the success of our system at reading a PDF will largely depend on the quality of the PDF! We use optical character recognition on image documents to parse the text, and generally, we have a high degree of success at creating accurate transcripts from images. The better the PDF, the more likely this is to work well. I hope that makes sense!
The PDF I have has very high quality text. Its a scanned document with lot of pages, but the LLMs still say “its empty”.
Thanks for surfacing this, and sorry for the frustration. I just filed PRD-715 so our engineers can dig into why OCR is handing us an empty transcript on scanned PDFs. If you can DM me that document (or a redacted sample) I can attach it straight to the investigation. I’ll circle back here as soon as we have an update.
Hi @hurmuli,
Thank you for flagging this.
We were able to reproduce the issue on our side, and it has been reported to our engineering team. They are currently investigating and working on a fix.
We will update you here as soon as a fix has been released.
Thank you again for bringing this to our attention.
1 Like
Hi @hurmuli,
Thank you again for flagging this for us.
Our engineering team has now fixed the issue. I tested it using the document from the session ID you provided, and it appears to be working as expected on our end.
When you have a chance, please try again and let us know if you notice anything unexpected.
Thank you again for helping us improve the platform.
1 Like