This document describes the features, key advantages, and use cases for education scenario recognition in Alibaba Cloud Optical Character Recognition (OCR). It also provides an API reference for the products.
Product overview
Optical Character Recognition (OCR) for education scenarios is tailored to recognize information such as test questions, mathematical formulas, and oral calculation problems. By iterating and optimizing high-precision, general-purpose OCR for educational contexts, this service provides recognition of text and mathematical formulas from images of math problems, detects and recognizes text in oral calculation exercises, and returns the bounding boxes and content of questions. This provides the foundational technology for smart education applications such as photo-based question search, whiteboard recognition, and automated grading. The service helps teachers with administrative tasks and accelerates the digital transformation of education.
Try our features: https://duguang.aliyun.com/experience?type=edu
Activate to get a free quota: https://ocr.console.aliyun.com/overview
Purchase page: https://common-buy.aliyun.com/?commodityCode=ocr_education_dp_cn#/buy
Features
Printed mathematical formula recognition
OCR recognizes printed mathematical formulas. Use cases include question entry, photo-based question search, and assignment grading.

Question recognition
OCR can effectively recognize questions in educational content. It improves recognition accuracy by tagging elements within the questions. Supported tag types include formulas, printed text, underlines, and images. This feature is a fundamental component for functions like photo-based question search.
For example, if you input an image of a multiple-choice math question that contains a radical expression, the recognition result panel will output the corresponding formula text, such as √(a²-1)=√(a-1)·√(a+1) and √(25a⁴)=5a².
Test paper question segmentation
OCR supports structured digital entry of test papers and supplementary materials across various subjects. It automatically segments questions, applies structured tags, and outputs structured content for the question, stem, options, and answer. This reduces manual data entry costs and is widely used to digitize teaching materials and grade test papers.
Test paper question segmentation provides the following API operations: Test Paper Question Segmentation – Scanned supports full-page question segmentation for scanned, printed K-12 materials in multiple layouts and across all subjects. It can automatically associate questions that span columns or pages, link questions with their corresponding images, and recognize formulas. Test Paper Question Segmentation – Photographed offers the same functionality as the scanned version but is optimized for camera-captured images. Fine-grained Structured Question Segmentation supports detailed layout and structured recognition, with capabilities including test paper structured recognition, teaching material structured recognition, multi-dimensional label definition, and formula recognition.
Oral calculation assessment
OCR recognizes and assesses elementary school oral calculation problems. It supports the four basic arithmetic operations with integers, mixed integer operations, comparisons, and finding the maximum or minimum numbers.

Full-page test paper recognition
This feature recognizes full pages of scanned K-12 materials across all subjects. The API recognizes printed text and formulas and returns their coordinates. It can also detect the location of images within questions and return their coordinates. This feature is suitable for full-page recognition and question retrieval from workbooks, teaching aids, and textbooks.

Fine-grained structured recognition
OCR supports structured recognition of teaching materials and test papers for multiple subjects. It automatically segments questions from a full page of a workbook, test paper, or teaching aid, and recognizes the text content and its coordinates. This is suitable for scenarios like fine-grained question production and intelligent test creation.
The fine-grained structured recognition API can return multi-dimensional labels such as title, topic, question, option, and answer. Its capabilities include test paper structured recognition, teaching material structured recognition, multi-dimensional label definition, and formula recognition.
Key advantages
High accuracy: Our models are trained on massive image samples, achieving industry-leading accuracy.
Low latency: Built on Alibaba's proprietary EAS online service cluster and highly optimized inference technology, the service delivers elastic, low-latency performance.
Advanced technology: Built on Alibaba Cloud's Platform for AI, the service uses a highly optimized PAI-TensorFlow deep learning framework to train state-of-the-art text detection and recognition models.
Stable service: We continuously deploy algorithm optimizations without impacting service stability.
Use cases
Question entry: Upload test paper images to automatically recognize question content. This improves transcription efficiency and reduces labor costs.
Answer search: Automatically recognize test questions with OCR and search for answers based on the results. This is widely used in educational software to assist with teaching.
Assignment grading: Recognize questions and automatically assess answers. The service can detect common typos, punctuation errors, and grammatical issues. This can significantly improve grading efficiency for teachers or be used in tutoring software to assist students.
API reference
Cloud Marketplace API (legacy) | Official website API (new) |