Showing posts with label Document Capture. Show all posts
Showing posts with label Document Capture. Show all posts

Tuesday, March 30, 2010

What is Advanced Data Extraction (ADE)?

OCR and Data Extraction

A big part of any OCR solution is the process of data capture and extraction.  Most document capture applications provide the ability to process the converted text and provide the extraction of expressions from the text.  So how can this  help?  Well, the ability to parse the OCR text provides automation, and allows you to populate fields based on what you find.

An example might be a form that has DOB: 1/2/1968

You want to extract everything to the right of DOB: from a document.  You can do this with an ADE engine.

Tuesday, February 9, 2010

OCR Software - Distributed vs. Centralized

OCR Software - Distributed vs. Centralized

Ah, the centralized versus distributed question...it is one that is continually asked in the scanning, capture and document capture space.  Most associate OCR Software with familiar desktop applications like eCopy Desktop, OmniPage, PaperPort, etc.  These provide, in a way, distribution of the overall OCR process to end users.

There are applications on the market that can provide centralized and controlled OCR capabilities, through either a server or a workstation deployment.  One example is PSI:Capture from PSIGEN, and advanced document capture application, that allows centralized OCR processing.  Why would you want to do this?  Well, in most cases, this type of OCR deplyment model is utilized in conjunction with a document capture system, for centralized capture, indexing, QA, OCR and migration to a centralized DM / ECM system.  Typically, these systems give a broad and expansive feature set, providing all different types of OCR functionality.

Saturday, January 30, 2010

OCR Software versus Document Capture Software

OCR Software versus Document Capture Software

So all OCR Software companies provide the ability to convert scanned files into text or searchable PDFs via the Optical Character Recognition process, but how do I capture/scan the images so the applications can do their conversion?

This is an interesting question.  Let's talk about Document Capture first.  This type of application is built from the ground up to scan/capture documents at a high rate of speed, provide the means to collect information about the documents through a number of means, and then export the document/data to a back end repository.  All document capture companies provide all types of OCR options, and usually OEM their OCR, ICR, OMR components from the major OCR application companies, like:  ABBYY, OpenText, Nuance, ReadSoft, etc.  Most of these companies have diversified their offering to include document capture, but their offerings far way short on the capture side in my opinion...they are OCR companies.

The real goal here is to get the best OCR results possible through a powerful OCR engine, and also minimize your time required to scan and process through the best document capture software.  So, if you are looking to do high volume OCR processing, I highly recommend choosing a capture application that utilizes your OCR engine of choice to get the best of both worlds.  I will write more on this topic in upcoming posts.  If you want some guidance on How to pick the right OCR Software, click on the link text.