Automated prompt optimization with samples

更新时间:
复制 MD 格式

The prompt feedback optimization feature uses the examples you provide to automatically generate an optimized prompt that produces your desired results.

Demonstration

Prompt before optimization

Prompt after optimization

The initial prompt is: Classify car-related articles into one of six categories: Product analysis, Car dealer sales, Classic nostalgia, Quality complaints, Sales performance, or Other. Output the result in JSON format as {"type":"<category_result>"}.

The optimized prompt includes the initial prompt (original classification instructions), added examples (few-shot examples) such as "ID3 sales article → Car dealer sales", and content hints (notes that clarify classification boundaries).

Incorrect inference result:

When given a text about the relationship between a car brand's quality complaints and sales, the model outputs the classification {"type":"Quality complaints"}. The task is completed using 265 input tokens and 6 output tokens.

Correct inference result:

The model correctly classifies the input text about the car brand, outputting {"type": "Sales performance"}. The status shows Execution completed. Statistics: 16 characters, 1,152 input tokens, and 7 output tokens.

Features

This feature does the following:

  1. Adds your sample data to the prompt.

  2. It automatically evaluates the prompt's results against your Evaluation Data over multiple rounds, reflecting on the outcomes to optimize the prompt and generate content hints.

Compared to automatic prompt optimization, this feature evaluates performance based on the data you provide. This produces higher-quality responses tailored to your use case.

Note

We recommend selecting Qwen-Max as the inference model.

Alibaba Cloud Model Studio automatically selects a portion of your sample data to add to the prompt. We recommend providing 5 to 10 entries, with at least one for each scenario.

Alibaba Cloud Model Studio evaluates the prompt based on its performance with the evaluation data and the selected inference model. For best results, we recommend providing at least 20 data entries. The more data you provide, the better the optimization.

image

Case study

Background

To improve content management efficiency on an automotive forum, you plan to use a large language model to classify articles. The classification criteria are as follows:

Which of the following categories does the car-related article below belong to:
"Product analysis",
"Car dealer sales",
"Classic nostalgia",
"Quality complaints",
"Sales performance",
"Other".
Please provide the final result in the JSON format: {"type":"<category_result>"}.

After using the preceding instructions as your prompt, you find the app fails to classify articles correctly. For example, an article that should be classified as "Sales performance" is miscategorized.

After some research, you realize that prompt engineering is the right solution. To improve the app's classification accuracy, you manually classify a set of typical articles. The following is your sample data:

Manually classified data

Sample data

query

answer

Title: Strange noise from a specific brand vehicle's chassis, owner claims it sounds like a "duck". Body: A video shows a car owner reporting a "duck-like" noise from their vehicle while driving. An inspection revealed the issue was a displaced lower control arm bushing. The owner had previously replaced the assembly with a non-original part at a repair shop, but the problem persisted due to a mismatched bore size. The repair team recommended replacing it with an original manufacturer assembly to ensure a proper fix and avoid recurring costs. They emphasized that skimping on repairs can lead to wasted time and money.

{"type": "Quality complaints"}

Title: A specific brand vehicle achieves top safety ratings, ensuring your peace of mind. Body: A specific brand vehicle achieved a full "excellent" rating in the China Insurance Automotive Safety Index (C-IASI) tests, demonstrating its superior safety performance. The vehicle, with its outstanding safety features and stable performance, has become a trusted choice for drivers. #SpecificBrand #SafetyRating. Video summary: A specific brand vehicle recently received a comprehensive excellent rating in rigorous safety tests. This achievement highlights the brand's leading position in safety performance. Its rich safety configuration and stable mechanical performance have earned high consumer recognition, providing drivers with all-around safety assurance and enhancing the brand's market competitiveness.

{"type": "Product analysis"}

Title: It's not every day you see a specific luxury model on the road. Body: This model is currently the most expensive sedan from its brand, yet it's seen less often than older, million-yuan models. I saw one yesterday at the car wash. At first, I mistook it for a different model, but the front seemed unusually wide. Then I saw the badging with letters below the logo on the trunk, confirming it was the luxury model. However, I just couldn't see what about its appearance justifies a price of 300,000-400,000 yuan. It looks less classy than the "Starry Sky" edition of another model. No wonder it sells so few units per year.

{"type": "Other"}

Title: Let's get straight to the point, there's a surprise below! Body: I admit, I can't get over you. Even though you don't answer my calls, reply to my messages, or agree to meet, I still can't let you go. Because you once said that when you need to buy a car, you'll contact me, and you'll even have your friends and family buy from me too. It's the end of the month, and I haven't sold a single unit of this specific model. The manager is on my case. This price of 10.xx is for one car only, no jokes. This is about my career. Drop a "1" in the comments, and I'll give you a rock-bottom price. #CarSales #GreatDeal #ForSale

{"type": "Car dealer sales"}

Title: Is a specific car model overstocked? Dealers: Clearing inventory, drive away for 15,000? #CarInventory #CarQuality. Video summary: At the end of the last century, a joint venture brand was a pioneer in the Chinese market, establishing a solid foundation with early models like the Santana and Jetta. These cars were not only early entrants but were also considered high-end, scarce commodities, symbolizing an improved standard of living. Although expensive at the time, cars have since become common transportation. Today, classic models from this brand, while much cheaper, are still seen on the road, reflecting enduring consumer recognition and the brand's historical influence.

{"type": "Classic nostalgia"}

Evaluation data

query

answer

Title: Joint venture between Brand X and Brand Y Group, "AutoTech," launches ISO9001 project. Body: In a specific month and year, a consulting firm initiated the ISO9001 quality management system project for "AutoTech," a joint venture established by Brand X and Brand Y Group. About AutoTech: Registered on a specific date with a capital of X billion yuan, AutoTech is a joint venture. Brand Y Group will invest approximately X billion euros in this partnership, its largest single investment in China in nearly X years. AutoTech will integrate Brand Y's hardware and software expertise with the software company's experience to develop full-stack advanced driver-assistance systems and autonomous driving solutions. These solutions will be deployed in Brand Y Group's electric vehicles in the Chinese market. About the consulting firm: The firm is a management consulting and training organization focused on sustainable development, smart manufacturing, and management systems. It has provided services to thousands of companies, including a significant percentage of top 500 domestic and global enterprises.

{"type": "Other"}

Title: A specific car model

{"type": "Other"}

Title: #CarModelA vs. #CarModelB# Depends on whether you prefer comfort or features. Body: #CarModelA vs. #CarModelB# Depends on whether you prefer comfort or features.

{"type": "Other"}

Title: After a week of careful research and comparison, I finally decided to buy the '24 model of Brand A's five-seater 2.0T 330. Its excellent performance and affordable price were the deciding factors. Body: The car's performance is impressive. The 2.0T engine provides ample power, making it fun to drive. The five-seat design meets my daily needs for both family outings and city commuting. The ride is comfortable, and handling is satisfactory. Price was a major factor. The car's sticker price was 135,000 yuan, and with insurance and taxes, the total was just over 150,000. In comparison, a car from Brand B was almost 50,000 yuan more expensive. Since both cars share the same platform and major components, the price advantage of Brand A was clear. Finally, regarding brand reputation, Brand A is well-regarded in Europe, though less known in this country. As part of a major automotive group, its quality and reliability are assured. I am very satisfied with my purchase. #CarBuyingExperience

{"type": "Product analysis"}

Title: Full body wrap for a specific brand's 2022 A7L 2.0TFSI 45TFSI S-line "White Mage" edition #A7L #A7 #SpecificBrand #PerformanceCar. Body: Full body wrap for a specific brand's 2022 A7L 2.0TFSI 45TFSI S-line "White Mage" edition #A7L #A7 #SpecificBrand #PerformanceCar.

{"type": "Other"}

You decide to use the prompt feedback optimization feature in Alibaba Cloud Model Studio to generate an improved prompt from your manually classified data.

Procedure

On the prompt > Feedback Optimization page in Alibaba Cloud Model Studio, click Create Optimization Task.

Step 1: Select an Inference Model. Alibaba Cloud Model Studio uses this model to perform multiple rounds of prompt evaluation.

Step 2: Enter the Original Prompt you want to optimize.

Just describe the goal of your task.

Step 3 (Optional): Select sample data. You can either upload a file directly or select data from a sample library.

The system adds the sample data to the optimized prompt. We recommend providing 5 to 10 data entries, with at least one entry for each scenario.

For this case study, we use the sample data in sample.xlsx.

Step 4: Upload Evaluation Data.

This data is used as the benchmark to evaluate the optimal prompt.

We recommend providing at least 20 data entries for your evaluation dataset. The more data you provide, the better the optimization results.

For this case study, we use the evaluation data in evaluation.xlsx.

Step 5: Start optimization.

Using the optimized prompt

  • You can save the optimized prompt as a prompt template or directly create an agent app based on it.