Scanned Image/PDF to Searchable Image/PDF

2019-05-29 21:34发布

Can anyone suggest me how to convert a scanned image into a searchable image or a scanned pdf to a searchable pdf ?
I have been stuck in this situation since quite a while now.
i have tried pdfocr application in ubuntu but no success.

标签： pdf-generation ocr

2条回答

我只想做你的唯一

2楼-- · 2019-05-29 22:03

Tesseract version 3.03 supports creation of searchable PDF from image. For PDF, you can use GhostScript to convert it to image before sending it to Tesseract.

https://github.com/tesseract-ocr/tesseract

0人赞添加讨论(0) 举报

孤傲高冷的网名

3楼-- · 2019-05-29 22:15

Currently, there is no right way of doing this on Ubuntu. All OCR engines output plain text and there is no way to add that text as a hidden layer on PDF over the image text.

Option 1: Use gscan2pdf which will make you a searchable PDF, but the OCRed text is placed in the top-left corner of the page, is invisible and much too small.

Option 2: Use PDF X-Change Viewer which has an option to OCR and works correctly by adding a text layer over the scanned image which is in concordance with it. You'll have to run it in wine, because it is a Windows application.

0人赞添加讨论(0) 举报

Scanned Image/PDF to Searchable Image/PDF

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间