Xerox M118i OmniPage SE User Guide - Page 22

What is optical character recognition, Omni SE’s OCR capabilities

Page 22 highlights

What is optical character recognition Optical character recognition is the process of extracting text from an image. This image can result from scanning a paper document or opening an electronic image file. Images do not have editable text characters; they have many tiny dots (pixels) that together form character shapes. These present a picture of the text on a page. During OCR, OmniPage SE analyzes the character shapes in an image and defines solutions to produce editable text. After OCR, you can save the resulting text to a variety of word-processing, desktop publishing or spreadsheet applications. OmniPage SE's OCR capabilities In addition to text recognition, OmniPage SE can retain the following elements of a document through the OCR process. Graphics Photos, logos, and drawings are examples of graphics. Text formatting Font types, sizes and styles (such as bold, italic and underlines) are examples of character formatting. Indents, tabs, margins and line spacing are examples of paragraph formatting. Page formatting Column structure, table formats, and placement of graphics and headings are examples of page formatting. The graphics, text and page formatting elements that OmniPage SE retains are determined by the settings you select. Refer to the Settings Guidelines in the online Help for more information about selecting settings. OmniPage SE only recognizes machine-generated characters such as offset or laserprinted or typewritten text. However, it can retain handwritten text, such as a signature, as a graphic. 22 Introduction

  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8
  • 9
  • 10
  • 11
  • 12
  • 13
  • 14
  • 15
  • 16
  • 17
  • 18
  • 19
  • 20
  • 21
  • 22
  • 23
  • 24
  • 25
  • 26
  • 27
  • 28
  • 29
  • 30
  • 31
  • 32
  • 33
  • 34
  • 35
  • 36
  • 37
  • 38
  • 39
  • 40
  • 41
  • 42
  • 43
  • 44
  • 45
  • 46
  • 47
  • 48
  • 49
  • 50
  • 51
  • 52
  • 53
  • 54
  • 55
  • 56
  • 57
  • 58
  • 59
  • 60
  • 61
  • 62
  • 63
  • 64
  • 65
  • 66
  • 67
  • 68
  • 69
  • 70
  • 71
  • 72
  • 73
  • 74
  • 75
  • 76
  • 77
  • 78
  • 79
  • 80
  • 81
  • 82
  • 83
  • 84
  • 85
  • 86
  • 87
  • 88
  • 89
  • 90
  • 91
  • 92
  • 93
  • 94
  • 95
  • 96
  • 97
  • 98
  • 99
  • 100
  • 101
  • 102

22
Introduction
What is optical character recognition
Optical character recognition is the process of extracting text from an
image. This image can result from scanning a paper document or
opening an electronic image file.
Images do not have editable text
characters; they have many tiny dots (pixels) that together form character
shapes. These present a picture of the text on a page.
During OCR, OmniPage SE analyzes the character shapes in an image
and defines solutions to produce editable text. After OCR, you can save
the resulting text to a variety of word-processing, desktop publishing or
spreadsheet applications.
OmniPage SE’s OCR capabilities
In addition to text recognition, OmniPage SE can retain the following
elements of a document through the OCR process.
Graphics
Photos, logos, and drawings are examples of graphics.
Text formatting
Font types, sizes and styles (such as
bold
,
italic
and underlines
) are
examples of character formatting. Indents, tabs, margins and line spacing
are examples of paragraph formatting.
Page formatting
Column structure, table formats, and placement of graphics and headings
are examples of page formatting.
The graphics, text and page formatting elements that OmniPage SE
retains are determined by the settings you select. Refer to the
Settings
Guidelines
in the online Help for more information about selecting
settings.
OmniPage SE only recognizes machine-generated characters such as offset or laser-
printed or typewritten text. However, it can retain handwritten text, such as a
signature, as a graphic.