So, Optical Character Recognition (OCR) is the process of recognizing computer generated text from an image, typically one that is scanned using document capture software. If you don't know the difference between OCR and Capture, see my other post here: OCR vs. Capture.
Intelligent Character Recognition (ICR) is the process of recognizing hand-printed or handwritten information from a scanned document. It utilizes the patterns of the pixels to match to specific written characters. This form of recognition is typically not as accurate as OCR, but there are several ways to make the accuracy acceptable, the main of which is to provide combed fields or spaced boxes to ensure character spacing, or whitespace between symbols.
Showing posts with label recognition software. Show all posts
Showing posts with label recognition software. Show all posts
Wednesday, March 9, 2011
Tuesday, January 26, 2010
Microsoft SharePoint and OCR
Microsoft SharePoint and OCR
Scanning with Microsoft SharePoint is an interesting endeavor, and typically the main reason for this undertaking is to have a searchable body of information. So what type of Optical Character Recognition (OCR) Software can be utilized with SharePoint? First of all, all the same rules apply in picking the right recognition software to do the conversion from image to text, as outlined in "How do I pick the right OCR Software?". You need to evaluate what you are trying to accomplish and look at your business process and workflows to get a good idea of how to initiate the conversion process. Below are some key questions when evaluating a SharePoint OCR Solution.
Are your paper images scanned en masse, through a centralized capture process?
If this is the case, you would typically do all of your OCR processing and recognition in front end document capture software. These application provide the fastest OCR engines, and their recognition processing time can be anywhere from 100-600 pages per minute, depending on the types of pages you are scanning.
Do you utilize MFPs / Copiers to scan document to sharepoint?
Most companies are trying to leverage their investment in their copier hardware to provide end users a great scanning and capture onramp to SharePoint. In this case, you typically want an OCR application that can provide recognition on the fly, and do the conversion process behind the scenes. Their are many MFP integrated applications on the market that can provide the OCR engine: iCapture, NSI AutoStore, eCopy to name a few.
Do the end users compile, combine and work with documents at their desktops?
In environments where end users are constantly working in their documents, and need desktop scanning access, typically and OCR Desktop application can be the best solution. These applications can put the control of the conversion process in the end user's hands, and can provide them OCR capability at the click of the mouse. Some apps in this class are eCopy Paperworks, PaperPort and OmniPage.
Do you want to SharePoint OCR PDFs?
Knowing what format and how you want to search can be critical, and having OCR PDFs in SharePoint can allow for full text search.
All of the OCR Solutions on this page focus on doing the process before the documents hit SharePoint. I will write an article later on solutions that can OCR documents within SharePoint Libraries later.
Scanning with Microsoft SharePoint is an interesting endeavor, and typically the main reason for this undertaking is to have a searchable body of information. So what type of Optical Character Recognition (OCR) Software can be utilized with SharePoint? First of all, all the same rules apply in picking the right recognition software to do the conversion from image to text, as outlined in "How do I pick the right OCR Software?". You need to evaluate what you are trying to accomplish and look at your business process and workflows to get a good idea of how to initiate the conversion process. Below are some key questions when evaluating a SharePoint OCR Solution.
Are your paper images scanned en masse, through a centralized capture process?
If this is the case, you would typically do all of your OCR processing and recognition in front end document capture software. These application provide the fastest OCR engines, and their recognition processing time can be anywhere from 100-600 pages per minute, depending on the types of pages you are scanning.
Do you utilize MFPs / Copiers to scan document to sharepoint?
Most companies are trying to leverage their investment in their copier hardware to provide end users a great scanning and capture onramp to SharePoint. In this case, you typically want an OCR application that can provide recognition on the fly, and do the conversion process behind the scenes. Their are many MFP integrated applications on the market that can provide the OCR engine: iCapture, NSI AutoStore, eCopy to name a few.
Do the end users compile, combine and work with documents at their desktops?
In environments where end users are constantly working in their documents, and need desktop scanning access, typically and OCR Desktop application can be the best solution. These applications can put the control of the conversion process in the end user's hands, and can provide them OCR capability at the click of the mouse. Some apps in this class are eCopy Paperworks, PaperPort and OmniPage.
Do you want to SharePoint OCR PDFs?
Knowing what format and how you want to search can be critical, and having OCR PDFs in SharePoint can allow for full text search.
All of the OCR Solutions on this page focus on doing the process before the documents hit SharePoint. I will write an article later on solutions that can OCR documents within SharePoint Libraries later.
Sunday, December 27, 2009
How do I pick the right OCR Software?
In the space of OCR Software, or Optical Character Recognition, it can be confusing to say the least on which option you should pick. It really comes down to the use case, or how you will utilize the software. Below are some great question to ask your self:
What do I need to convert with my OCR Software?
This question is very important, and it really comes down to what you are looking to output with your software. Do you want a word file that you can edit, or are you just looking to create a searchable PDF? Many engines are tuned for accuracy, and will give you the best formatted output, others are built for speed. Omni-page is an excellent engine for creating nicely formatted output, but can be rather slow due to its focus on acuracy. A production engine, like PSI:Capture, which offers multiple OCR choices, can give you great flebility, no matter your ouput choice.
Are they pre-existing images, or ones that I will scan? PDFs or TIFFs?
It is really important when you are choosing Optical Character Recognition Software, to make sure that you have all the functionality you require, whether you are scanning, or just processing non-searchable PDFs from a directory. Most of the OCR Software will let you choose the file that you perform recognition on, and others will let you scan in paper for conversion. If you are utilizing MFPs or Scanning copiers, and want to perform OCR on the scanned documents, you may want to choose a product that performs auto-import, or one that is focused on MFP Scanning. Also, you want flexibility in the types of file you can process, and want to be able to OCR any image type: PDF, TIFF, JPG, GIF, BMP, etc.
How fast can I do conversions?
So, some engines are built for OCR Accuracy, others built for speed in the OCR process. Most of the desktop engines, like eCopy Desktop, provide a good mix of both. Other engines, like Glyphreader or Docustar, provide the ability to choose whether you want speed or accuracy in your OCR results. It is always good to choose a document capture option that allows you multiple OCR engine options to perform diffferent recognition tasks.
How ddo I get the best accuracy in the OCR ouput?
All of the OCR Software mentioned within this post reuires a high quality image for the best recognition accuracy. With that said, a high quality scanning software with image processing options will lead to the best OCR accuracy when converting from image to text. So what does image processing have to do with OCR Software? The cleaner the image, the better the accuracy, and if you can deskew, despeckle, deshade and sharpen text, you will get better OCR results.
What do I need to convert with my OCR Software?
This question is very important, and it really comes down to what you are looking to output with your software. Do you want a word file that you can edit, or are you just looking to create a searchable PDF? Many engines are tuned for accuracy, and will give you the best formatted output, others are built for speed. Omni-page is an excellent engine for creating nicely formatted output, but can be rather slow due to its focus on acuracy. A production engine, like PSI:Capture, which offers multiple OCR choices, can give you great flebility, no matter your ouput choice.
Are they pre-existing images, or ones that I will scan? PDFs or TIFFs?
It is really important when you are choosing Optical Character Recognition Software, to make sure that you have all the functionality you require, whether you are scanning, or just processing non-searchable PDFs from a directory. Most of the OCR Software will let you choose the file that you perform recognition on, and others will let you scan in paper for conversion. If you are utilizing MFPs or Scanning copiers, and want to perform OCR on the scanned documents, you may want to choose a product that performs auto-import, or one that is focused on MFP Scanning. Also, you want flexibility in the types of file you can process, and want to be able to OCR any image type: PDF, TIFF, JPG, GIF, BMP, etc.
How fast can I do conversions?
So, some engines are built for OCR Accuracy, others built for speed in the OCR process. Most of the desktop engines, like eCopy Desktop, provide a good mix of both. Other engines, like Glyphreader or Docustar, provide the ability to choose whether you want speed or accuracy in your OCR results. It is always good to choose a document capture option that allows you multiple OCR engine options to perform diffferent recognition tasks.
How ddo I get the best accuracy in the OCR ouput?
All of the OCR Software mentioned within this post reuires a high quality image for the best recognition accuracy. With that said, a high quality scanning software with image processing options will lead to the best OCR accuracy when converting from image to text. So what does image processing have to do with OCR Software? The cleaner the image, the better the accuracy, and if you can deskew, despeckle, deshade and sharpen text, you will get better OCR results.
Subscribe to:
Posts (Atom)