OCR Self-learning overview
As part of our ongoing service improvements, OCR Self-learning is no longer recommended. This product will not receive further optimizations or maintenance, and its APIs will be gradually deprecated and discontinued.
To ensure business continuity, we recommend migrating to the Model Studio vision large model solution. For an example of a large model solution, see Text Extraction.
Migration complexity: The vision large model solution replaces the legacy APIs. We provide detailed documentation and sample code to simplify the migration.
Feature overview
OCR Self-learning is an all-in-one tool platform designed for enterprise and individual developers with no prior algorithm experience. It provides a fully visual workflow that allows you to configure templates, process and label data, build and train models, and manage deployment. The platform uses leading-edge AI technologies such as few-shot learning, intelligent pre-labeling, and joint visual-semantic learning. This lets you digitize documents and extract information for your specific use cases at a low cost.
Customization tools: Use self-service tools to customize models for your business scenarios and build data-driven AI services.
Multimodal information extraction: Extract custom multimodal information through a reliable and easy-to-use service.
Few-shot cold start: Supports service customization with as little as a single image.
High-efficiency customization: Build and customize an end-to-end AI model in hours, significantly reducing business wait times.
User-friendly interaction: A visual interface simplifies model training and usage for all users.
Feature details
The OCR Self-learning platform supports self-service training for two main project types: templates and models. You can configure a template or label a small amount of data to train a custom AI model.
Template:
Custom KV template
Configure a template by using a single sample image, defining fields and rules. No additional image labeling or training time is required. This lets you immediately extract custom fields from fixed-layout documents and tickets. For more information, see the Operation Guide.
Custom table template
Configure a template by using a single sample table image, defining fields and rules. No extra labeling or training is needed. This lets you extract custom cells from single-page tables that have a fixed layout and visible borders. For more information, see the Operation Guide.
Model:
Document and ticket information extraction
Extract key fields from documents, certificates, and vouchers with relatively fixed layouts. This data-driven approach involves labeling a small data sample to train a model. For more information, see the Operation Guide.
Table information extraction
Extract key fields from tables and forms with relatively fixed layouts. This data-driven approach involves labeling a small data sample to train a model. For more information, see the Operation Guide.
Long document information extraction
Extract key information from multi-layout, unstructured long documents. This data-driven approach involves labeling a small data sample to train a model. For more information, see the Operation Guide.
Toolbox:
Classifier management
Add keywords and classification data to associate different template or model types with a classifier. This lets a single API endpoint accept multiple sample data types, route requests to the correct capability, and perform information extraction.
Field type management
Configure field types, primarily for fields with common business or industry-specific properties. You can use this feature to correct field errors to improve recognition accuracy or to normalize data.
When should I use a custom template versus an information extraction model?
Custom template: Configure with just a single sample image, with no model training required. This is ideal for quickly validating a business use case during the cold-start phase, especially for fixed-layout data where high accuracy is not critical.
Information extraction model: Follows a standard "label data, train model" workflow. Use the visual tools to build a custom model for your business. This is suitable for stable projects with sufficient sample data that require high accuracy. It works best when data layouts are relatively fixed or enumerable.
Value proposition
Turn your data into assets:
Manage your data assets from end to end, including uploading, processing, and labeling. With one-stop preprocessing and labeling tools and a guided visual interface, users with no algorithm experience can create and publish a custom template in under 5 minutes. This helps you continuously build up your data assets and drive business transformation.
Turn your models into business services:
Leverage built-in, general-purpose multimodal AI capabilities and your accumulated data assets to train a custom model with a single click. By using core technologies such as pre-trained models, multimodal vision-language algorithms, and few-shot information extraction, you can meet your business needs with greater efficiency and accuracy.
Centralize management on a single platform:
Use an all-in-one tool platform with management tools for the entire workflow, from data asset management to model building, training, and deployment. You can continuously track model performance and business impact. In the future, you can feed positive and negative samples back into the system to enable lifelong learning and continuous iteration, creating a closed loop that consistently enhances business value.
Product advantages
Multimodal document information extraction
The platform specializes in visual document information extraction, simplifying personalized data extraction from complex visual documents. It provides stable service, accurate results, and intelligent workflows.
No-code customization
Techniques like few-shot learning simplify model training. This allows users with no algorithm background to use their own data to build custom models, transforming data assets into service assets.
High-accuracy models
The platform includes built-in, large-scale multimodal pre-trained models, high-accuracy text recognition models for various scenarios, and a unified information extraction model. This ensures the accuracy required for no-code modeling across different use cases.
Efficient model production
Built-in intelligent pre-labeling and an easy-to-use, all-in-one labeling suite dramatically increase labeling efficiency. The included base pre-trained models also significantly speed up the training process during the fine-tuning phase.
Flexible deployment modes
Supports a high-availability public cloud deployment mode and an on-premises private deployment to meet different customer needs.
Use cases
Invoice and document extraction
Extract key-value information from various documents and tickets with an accuracy of up to 95%. It is suitable for scenarios where layouts are relatively fixed and enumerable.
Table and form parsing
Extract information from various tables and forms with an accuracy of up to 95%. It is ideal for use cases with relatively fixed and enumerable layouts.
Unstructured long document parsing
Automate information extraction from various unstructured documents with an accuracy of up to 85%. This is suitable for processing multi-page, unstructured documents.
Announcement and official document processing
Use OCR Self-learning to extract information from documents with non-fixed layouts and styles, such as announcements and official papers.
Contact us
Contact us through our DingTalk groups.
[Official] Alibaba Cloud OCR Self-learning User Q&A Group: 26560014923
[Official] Alibaba Cloud OCR Public Cloud Customer Group: 35208328
[Official] Alibaba Cloud Document Intelligence Customer Group: 44854217