Feature introduction
OCR Document Self-Learning is a one-stop platform for developers and businesses, even those with no algorithm experience. It provides a visual, end-to-end workflow to configure templates, process and annotate data, build and train models, and deploy services. The platform uses advanced Artificial Intelligence (AI) technologies, such as few-shot training, intelligent pre-annotation, and visual-semantic joint learning. This enables you to digitize and automate document workflows for custom scenarios at a low cost.
Provides customizable tools to build custom models for your business scenarios. You can create AI services driven by your own data.
Multi-modal information extraction. You can extract custom multi-modal information to create useful and reliable services.
Supports few-shot cold starts. You can customize a service with as little as a single image.
Improves customization efficiency. You can customize an end-to-end AI model in hours, which greatly reduces business wait times.
User-friendly interaction. Lowers the barrier to model training with a visual and interactive interface.
Feature details
The OCR Document Self-Learning platform supports self-service training for two project types: templates and models. You can configure templates or annotate a small amount of data to train AI models that are better suited for your business scenarios.
Templates:
Custom key-value template
You can configure a template image with field information and rules to extract custom fields from fixed-layout documents and certificates without annotating more images or waiting for training. For more information, see the User Guide.
Custom table template
You can configure a template table image with field information and rules to extract custom cells from single-page, fixed-layout tables with borders, without annotating more images or waiting for training. For more information, see the User Guide.
Models:
Document and certificate information extraction
This feature is data-driven. You can annotate and train a small data sample to extract key fields from documents, certificates, and vouchers with relatively fixed layouts. For more information, see the User Guide.
Table information extraction
This feature is data-driven. You can annotate and train a small data sample to extract key fields from tables and forms with relatively fixed layouts. For more information, see the User Guide.
Long document information extraction
This feature is data-driven. You can annotate and train a small data sample to extract key information from unstructured long documents with multiple layouts. For more information, see the User Guide.
Toolbox:
Classifier management
You can add keywords and categorical data to associate different template or model types with a single classifier. This allows a single API endpoint to accept multiple types of sample data and route them to the correct service for information extraction.
Field type management
You can configure field types. This feature is mainly for fields with common business or industry properties. You can use it for field correction to improve recognition accuracy or for data normalization.
Choosing between custom templates and information extraction models
Custom templates: You can configure a template with just one sample image. No model training is required. This method is suitable for the rapid validation phase of a business cold start, where data layouts are fixed and high accuracy for field extraction is not required.
Information extraction models: These models require a standard data annotation and model training workflow. You can use visual tools to train and customize a model for your business. This method is suitable for the stable phase of a business, where data layouts are relatively fixed or enumerable, you have sufficient sample data, and you require high accuracy for information extraction.
Value proposition
Turn data into assets:
You can manage the entire lifecycle of your data assets, from uploading and processing to annotation. The platform provides one-stop tools for pre-processing and annotation. With visual guidance, users who have no algorithm background can create and publish a custom template in less than 5 minutes. This process helps you build your data assets and drive business upgrades.
Turn models into business services:
You can use pre-built multi-modal AI capabilities and your data assets to train custom models with a single click. Core technologies, such as pre-trained models, multi-modal algorithms, and few-shot information extraction, help you meet your business needs with greater efficiency and accuracy.
Centralize management on one platform:
You can use the one-stop platform for end-to-end management of data assets, model building, training, and deployment. You can also track model evaluations and business performance. Using continuous feedback from business samples, you can enable lifelong learning and iterative upgrades for your models. This creates a closed-loop system that continuously adds value to your business.
Benefits
Multi-modal document information extraction
The platform focuses on visual document information extraction and addresses the challenges of extracting custom information from complex visual documents. The result is a stable, accurate, and intelligent self-learning platform.
No-code self-service customization
Technologies such as few-shot learning lower the barrier to model training. This allows users who have no algorithm background to use their own data to customize models and turn data assets into service assets.
High-precision model performance
The platform includes built-in, large-scale multi-modal pre-trained models, high-precision text recognition models for multiple scenarios, and a unified information extraction model. These features meet the accuracy requirements for no-code modeling in different scenarios.
Efficient model production
Built-in intelligent pre-annotation and an easy-to-use, one-stop annotation suite greatly improve annotation efficiency. The built-in base pre-trained model significantly improves training efficiency during the fine-tuning phase.
Flexible deployments
The platform supports high-availability public cloud and on-premises private deployments to meet the needs of different customers.
Scenarios
Receipt and document extraction
You can extract key-value (KV) information from various documents and receipts with an accuracy rate of up to 95%. This feature is suitable for scenarios with relatively fixed and enumerable layouts.
Table and form parsing
You can extract information from various tables and forms with an accuracy rate of up to 95%. This feature is suitable for scenarios with relatively fixed and enumerable layouts.
Unstructured long document parsing
You can automatically extract information from various unstructured documents with an accuracy rate of up to 85%. This feature is suitable for processing unstructured, multi-page documents.
Announcement and official document processing
You can extract information from documents such as announcements and official papers. Use the Document Self-Learning platform to process documents with non-fixed layouts and styles.
Contact us
If you have any questions, contact us through our DingTalk groups.
[Official] Alibaba Cloud OCR Document Self-Learning User Q&A Group: 26560014923
[Official] Alibaba Cloud OCR Public Cloud Customer Group: 35208328
[Official] Alibaba Cloud Document Mind Customer Group: 44854217