《OCR 工具与项目实战》
OCR(光学字符识别)用于从图像中提取文字,场景覆盖扫描件、车牌、票据等。本文整理常用工具(Tesseract、open-ocr)与深度学习方法的学习资源。
1 Tesseract(开源 OCR 引擎)
Tesseract 是最常用的开源文本识别引擎。
1.1 安装
# macOS
brew install tesseract tesseract-lang
# Ubuntu / Debian
sudo apt install tesseract-ocr tesseract-ocr-chi-sim # chi-sim 为简体中文语言包
1.2 命令行使用
# 识别图片中的文字,输出到终端
tesseract image.png stdout
# 输出到文件,默认英文
tesseract image.png out
# 指定中文 + 英文
tesseract image.png out -l chi_sim+eng
# 保留版面/表格布局
tesseract image.png out --psm 6
常用 --psm(页面分割模式):
| 值 | 含义 |
|---|---|
| 6 | 假设一个均匀文本块 |
| 3 | 自动分页(默认) |
| 11 | 稀疏文本,无特定顺序 |
| 7 | 单行文字 |
1.3 Python 调用
pip install pytesseract pillow
from PIL import Image
import pytesseract
text = pytesseract.image_to_string(Image.open("image.png"), lang="eng")
print(text)
中文识别前建议先对图像做灰度化、二值化、去噪,能显著提升正确率。
2 open-ocr 项目
Tesseract 的容器化服务封装,便于以 HTTP API 方式提供 OCR 能力:
# docker-compose 一键起服务(示意)
services:
openocr:
image: tleyden5iwx/open-ocr
command: ["/usr/local/bin/run_service.sh", "py-tesseract-ocr"]
ports:
- "9292:9292"
curl -F "image=@image.png" http://localhost:9292/ocr
3 深度学习 OCR 思路
小数据集上即可上手深度学习 OCR,常用路线:
- CNN → CTC
- 参考用 Keras + Supervisely 15 分钟训练 OCR:https://hackernoon.com/latest-deep-learning-ocr-with-keras-and-supervisely-in-15-minutes-34aecd630ed8
- EMNIST(扩展手写字母数字集)替代 MNIST 做字符分类。
- 目标检测 + 识别两阶段
- 先用检测模型(YOLO / SSD)定位文字区域,再送入识别模型。
- 车牌检测参考:https://towardsdatascience.com/number-plate-detection-with-supervisely-and-tensorflow-part-1-e84c74d4382c
- 端到端
- 成熟方案如 PaddleOCR、EasyOCR,开箱即用且支持中文。
4 附:相关教程
- OpenCV 图像读取与显示入门:https://opencv-python-tutroals.readthedocs.io/en/latest/py_tutorials/py_gui/py_image_display/py_image_display.html#display-image
- 用 CNTK 从零构建 OCR(MNIST 之外的进阶思路):https://blogs.msdn.microsoft.com/uk_faculty_connection/2017/11/20/bored-of-mnist-lets-build-your-own-ocr-deep-learning-computer-vision-ai-using-microsoft-cntk-with-emnist-step-by-step-guide/
5 常见问题
- 中文识别乱码:需安装并指定
chi_sim/chi_sim+eng语言包。 - 识别率低:优先做图像预处理(缩放、二值化)与
--psm调参。 - 大批量处理:用 Tesseract 的 batch 接口或接入 open-ocr / PaddleOCR 服务化。
阅读 —
·
全站 —