OCR Self-learning overview

Updated at:
Important
  • As part of our ongoing service improvements, OCR Self-learning is no longer recommended. This product will not receive further optimizations or maintenance, and its APIs will be gradually deprecated and discontinued.

  • To ensure business continuity, we recommend migrating to the Model Studio vision large model solution. For an example of a large model solution, see Text Extraction.

  • Migration complexity: The vision large model solution replaces the legacy APIs. We provide detailed documentation and sample code to simplify the migration.

Feature overview

OCR Self-learning is an all-in-one tool platform designed for enterprise and individual developers with no prior algorithm experience. It provides a fully visual workflow that allows you to configure templates, process and label data, build and train models, and manage deployment. The platform uses leading-edge AI technologies such as few-shot learning, intelligent pre-labeling, and joint visual-semantic learning. This lets you digitize documents and extract information for your specific use cases at a low cost.

  • Customization tools: Use self-service tools to customize models for your business scenarios and build data-driven AI services.

  • Multimodal information extraction: Extract custom multimodal information through a reliable and easy-to-use service.

  • Few-shot cold start: Supports service customization with as little as a single image.

  • High-efficiency customization: Build and customize an end-to-end AI model in hours, significantly reducing business wait times.

  • User-friendly interaction: A visual interface simplifies model training and usage for all users.

Feature details

The OCR Self-learning platform supports self-service training for two main project types: templates and models. You can configure a template or label a small amount of data to train a custom AI model.

Template:

  • Custom KV template

    • Configure a template by using a single sample image, defining fields and rules. No additional image labeling or training time is required. This lets you immediately extract custom fields from fixed-layout documents and tickets. For more information, see the Operation Guide.

  • Custom table template

    • Configure a template by using a single sample table image, defining fields and rules. No extra labeling or training is needed. This lets you extract custom cells from single-page tables that have a fixed layout and visible borders. For more information, see the Operation Guide.

Model:

Toolbox:

  • Classifier management

    • Add keywords and classification data to associate different template or model types with a classifier. This lets a single API endpoint accept multiple sample data types, route requests to the correct capability, and perform information extraction.

  • Field type management

    • Configure field types, primarily for fields with common business or industry-specific properties. You can use this feature to correct field errors to improve recognition accuracy or to normalize data.

Note

When should I use a custom template versus an information extraction model?

Custom template: Configure with just a single sample image, with no model training required. This is ideal for quickly validating a business use case during the cold-start phase, especially for fixed-layout data where high accuracy is not critical.

Information extraction model: Follows a standard "label data, train model" workflow. Use the visual tools to build a custom model for your business. This is suitable for stable projects with sufficient sample data that require high accuracy. It works best when data layouts are relatively fixed or enumerable.

Value proposition

  • Turn your data into assets:

    • Manage your data assets from end to end, including uploading, processing, and labeling. With one-stop preprocessing and labeling tools and a guided visual interface, users with no algorithm experience can create and publish a custom template in under 5 minutes. This helps you continuously build up your data assets and drive business transformation.

  • Turn your models into business services:

    • Leverage built-in, general-purpose multimodal AI capabilities and your accumulated data assets to train a custom model with a single click. By using core technologies such as pre-trained models, multimodal vision-language algorithms, and few-shot information extraction, you can meet your business needs with greater efficiency and accuracy.

  • Centralize management on a single platform:

    • Use an all-in-one tool platform with management tools for the entire workflow, from data asset management to model building, training, and deployment. You can continuously track model performance and business impact. In the future, you can feed positive and negative samples back into the system to enable lifelong learning and continuous iteration, creating a closed loop that consistently enhances business value.

Product advantages

Multimodal document information extraction

The platform specializes in visual document information extraction, simplifying personalized data extraction from complex visual documents. It provides stable service, accurate results, and intelligent workflows.

No-code customization

Techniques like few-shot learning simplify model training. This allows users with no algorithm background to use their own data to build custom models, transforming data assets into service assets.

High-accuracy models

The platform includes built-in, large-scale multimodal pre-trained models, high-accuracy text recognition models for various scenarios, and a unified information extraction model. This ensures the accuracy required for no-code modeling across different use cases.

Efficient model production

Built-in intelligent pre-labeling and an easy-to-use, all-in-one labeling suite dramatically increase labeling efficiency. The included base pre-trained models also significantly speed up the training process during the fine-tuning phase.

Flexible deployment modes

Supports a high-availability public cloud deployment mode and an on-premises private deployment to meet different customer needs.

Use cases

Invoice and document extraction

Extract key-value information from various documents and tickets with an accuracy of up to 95%. It is suitable for scenarios where layouts are relatively fixed and enumerable.

Table and form parsing

Extract information from various tables and forms with an accuracy of up to 95%. It is ideal for use cases with relatively fixed and enumerable layouts.

Unstructured long document parsing

Automate information extraction from various unstructured documents with an accuracy of up to 85%. This is suitable for processing multi-page, unstructured documents.

Announcement and official document processing

Use OCR Self-learning to extract information from documents with non-fixed layouts and styles, such as announcements and official papers.

Contact us

Contact us through our DingTalk groups.

[Official] Alibaba Cloud OCR Self-learning User Q&A Group: 26560014923

[Official] Alibaba Cloud OCR Public Cloud Customer Group: 35208328

[Official] Alibaba Cloud Document Intelligence Customer Group: 44854217