Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Ideally you'd get something besides a PDF, which is a pain in the neck to turn into a 'real' eBook.


I'll probably OCR the PDF and see where I can go from there. It's a big jump start that I couldn't have easily done on my own.


The 1DollarScan website says they do the OCR for you. They also claim that the result is searchable, which of course it couldn't be if they didn't OCR it.


I may still want to OCR it myself for more control over the process (particularly in identifying and correcting OCR errors). In my experience, there's a difference in expectation of quality between OCR to make a PDF searchable and OCR to generate a standalone text file.


Why is that? .PDF isn't proprietary enough?


PDF is really a display format, whereas epub/mobi are more HTMLish in that they are not quite so specific in how they want things displayed, and thus can be flowed into different screen sizes pretty easily.


PDF is not proprietary.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: