Showing posts with label OCR results. Show all posts
Showing posts with label OCR results. Show all posts
Wednesday, May 9, 2012
OCR and PC Architecture
So just how important is your PC hardware when looking to use OCR Software? Many of the desktop products do not take advantage of multi-core CPUs, and can have laggard performance numbers when it comes to Optical Character Recognition, Intelligent Character Recognition and Optical Mark Recognition. Currently, playing with PSI:Capture, which offers a number of OCR options, and they have single, dual and quad core enablement in their licensing. Dual core runs about 1.7 times the speed, and quad core gives a 2.7x improvement.
Friday, February 19, 2010
Are index fileds really necessary when you have the full text OCR?
Full text OCR
Ah, the old debate, do I just perform optical character recognition on all my scanned documents, make them searchable OCR PDFs, and rely on the OCR to retrieve documents? Why use index fields when I already have all the converted text?
Index fields, or performing the indexing process, provides structured data about the documents. This data can be utilized, especially when using document capture software, to link into columns and index fields in your document management system. Index fields provide faster retrieval, especially if you want to be able to retrieve through specifying several criteria. Relying on OCR, or the recognized text can get you in trouble. First of all, you are assuming that the document will alwyas have recognized text, and that all the items that you are searching for are in the text. Secondly, depdning on the type of OCR format you have, you may have to just find the document, and then open and parse what you are looking for. This can also lead to false positives in retrieval if many documents have the same terms in their OCR text.
Ah, the old debate, do I just perform optical character recognition on all my scanned documents, make them searchable OCR PDFs, and rely on the OCR to retrieve documents? Why use index fields when I already have all the converted text?
Index fields, or performing the indexing process, provides structured data about the documents. This data can be utilized, especially when using document capture software, to link into columns and index fields in your document management system. Index fields provide faster retrieval, especially if you want to be able to retrieve through specifying several criteria. Relying on OCR, or the recognized text can get you in trouble. First of all, you are assuming that the document will alwyas have recognized text, and that all the items that you are searching for are in the text. Secondly, depdning on the type of OCR format you have, you may have to just find the document, and then open and parse what you are looking for. This can also lead to false positives in retrieval if many documents have the same terms in their OCR text.
Labels:
centralized ocr,
index,
ocr performance,
OCR results
Sunday, December 27, 2009
How do I pick the right OCR Software?
In the space of OCR Software, or Optical Character Recognition, it can be confusing to say the least on which option you should pick. It really comes down to the use case, or how you will utilize the software. Below are some great question to ask your self:
What do I need to convert with my OCR Software?
This question is very important, and it really comes down to what you are looking to output with your software. Do you want a word file that you can edit, or are you just looking to create a searchable PDF? Many engines are tuned for accuracy, and will give you the best formatted output, others are built for speed. Omni-page is an excellent engine for creating nicely formatted output, but can be rather slow due to its focus on acuracy. A production engine, like PSI:Capture, which offers multiple OCR choices, can give you great flebility, no matter your ouput choice.
Are they pre-existing images, or ones that I will scan? PDFs or TIFFs?
It is really important when you are choosing Optical Character Recognition Software, to make sure that you have all the functionality you require, whether you are scanning, or just processing non-searchable PDFs from a directory. Most of the OCR Software will let you choose the file that you perform recognition on, and others will let you scan in paper for conversion. If you are utilizing MFPs or Scanning copiers, and want to perform OCR on the scanned documents, you may want to choose a product that performs auto-import, or one that is focused on MFP Scanning. Also, you want flexibility in the types of file you can process, and want to be able to OCR any image type: PDF, TIFF, JPG, GIF, BMP, etc.
How fast can I do conversions?
So, some engines are built for OCR Accuracy, others built for speed in the OCR process. Most of the desktop engines, like eCopy Desktop, provide a good mix of both. Other engines, like Glyphreader or Docustar, provide the ability to choose whether you want speed or accuracy in your OCR results. It is always good to choose a document capture option that allows you multiple OCR engine options to perform diffferent recognition tasks.
How ddo I get the best accuracy in the OCR ouput?
All of the OCR Software mentioned within this post reuires a high quality image for the best recognition accuracy. With that said, a high quality scanning software with image processing options will lead to the best OCR accuracy when converting from image to text. So what does image processing have to do with OCR Software? The cleaner the image, the better the accuracy, and if you can deskew, despeckle, deshade and sharpen text, you will get better OCR results.
What do I need to convert with my OCR Software?
This question is very important, and it really comes down to what you are looking to output with your software. Do you want a word file that you can edit, or are you just looking to create a searchable PDF? Many engines are tuned for accuracy, and will give you the best formatted output, others are built for speed. Omni-page is an excellent engine for creating nicely formatted output, but can be rather slow due to its focus on acuracy. A production engine, like PSI:Capture, which offers multiple OCR choices, can give you great flebility, no matter your ouput choice.
Are they pre-existing images, or ones that I will scan? PDFs or TIFFs?
It is really important when you are choosing Optical Character Recognition Software, to make sure that you have all the functionality you require, whether you are scanning, or just processing non-searchable PDFs from a directory. Most of the OCR Software will let you choose the file that you perform recognition on, and others will let you scan in paper for conversion. If you are utilizing MFPs or Scanning copiers, and want to perform OCR on the scanned documents, you may want to choose a product that performs auto-import, or one that is focused on MFP Scanning. Also, you want flexibility in the types of file you can process, and want to be able to OCR any image type: PDF, TIFF, JPG, GIF, BMP, etc.
How fast can I do conversions?
So, some engines are built for OCR Accuracy, others built for speed in the OCR process. Most of the desktop engines, like eCopy Desktop, provide a good mix of both. Other engines, like Glyphreader or Docustar, provide the ability to choose whether you want speed or accuracy in your OCR results. It is always good to choose a document capture option that allows you multiple OCR engine options to perform diffferent recognition tasks.
How ddo I get the best accuracy in the OCR ouput?
All of the OCR Software mentioned within this post reuires a high quality image for the best recognition accuracy. With that said, a high quality scanning software with image processing options will lead to the best OCR accuracy when converting from image to text. So what does image processing have to do with OCR Software? The cleaner the image, the better the accuracy, and if you can deskew, despeckle, deshade and sharpen text, you will get better OCR results.
Subscribe to:
Posts (Atom)