Text-to-image

更新时间:
复制 MD 格式

Generate images from text descriptions with the text-to-image API. This service, provided by Alibaba Cloud Model Studio, features the Wan, Qwen-Image, and Z-Image model families.

Try it online: Beijing | Singapore

Model performance

Qwen-Image

Complex layout

c13672c7-6e05-9aeb-8d2c-8c4b66af5861_qwen_image3_serving_output_0

Long paragraph

594fe793-69a8-9ce0-bef0-1e36958f882a_qwen_image3_serving_output_0

Realistic portrait

2beecb7f-c3c3-9c10-9458-04325512bb8a_qwen_image3_serving_output_0

UI design

image (92)

PPT

image (90)

Illustration design

e38cc145-78f6-92a7-9bdb-141be17dba04_qwen_image3_serving_output_0

Prompts

Complex layout: Create a 16:9 wide horizontal e-commerce storefront hero banner infographic for "TodoTi," a Mexican-inspired colorful lifestyle brand. The design should look like a professional online shop header for desktop and tablet, with clear visual hierarchy, festive marketing energy, and readable English typography. Background: vibrant Mexican folk color palette of bright yellow, warm red, peacock blue, and emerald green, blended with soft watercolor washes, subtle handmade paper texture, and fine grain. The top and bottom edges are decorated with traditional papel picado cut-paper banners made of repeating diamonds, triangles, and wave shapes in alternating yellow, red, green, and purple. The papel picado should look light and gently fluttering, with cut-out areas revealing a deep starry-night blue background. Below the top papel picado and above the bottom papel picado, add a continuous Aztec geometric border of stepped lines and diamond motifs in white and antique gold. Four corners: add small 3D decorative Talavera ceramic pots with blue-and-white floral patterns, glossy glaze highlights, and miniature potted cacti with rich green texture and fine spines. These corner elements should balance the composition and add handcrafted detail. Left third: brand identity zone. Feature a large 3D wordmark "TodoTi." The letter "T" extends into two green cactus pads with clear vein texture; the letter "o" contains a red-and-yellow sunburst with sixteen thin rays; the curve of the letter "d" is decorated with blue-and-white Talavera ceramic patterns featuring three peonies and winding vines; the dot of the letter "i" is a glossy red cherry tomato with a white highlight. The wordmark has bold black outlines, high-saturation gradient fills, and a soft drop shadow. Below the wordmark, place the slogan: "TodoTi — Your Colorful Lifestyle Hub." The slogan is in rounded bold white sans-serif with a subtle red glow, centered and highly legible. Below it, add a smaller handwritten-style subtitle in bright yellow: "Mexican-inspired warmth to brighten every day." Behind the wordmark, add a faint semi-circular Aztec sun calendar relief in antique gold and bronze, with fine geometric patterns and a stylized face motif featuring diamond eyes and a rectangular mouth. The relief should glow softly with warm golden light. On the left edge, stack three small circular badges with peacock-blue backgrounds, white text, and stitched dashed borders: "Curated Picks," "Direct Sourcing," and "Thoughtful Service." Center-left: a cheerful Mexican folk-art lifestyle illustration. A cute Catrina skeleton girl wears a wide straw sombrero with red-and-green woven tassels and a colorful embroidered poncho in warm red and peacock blue. Her poncho is decorated with Aztec geometric patterns and blooming marigolds with radiating petals. Her face has elegant Day of the Dead makeup: black heart-shaped eye details with white dashed outlines, red rose cheek motifs, and a warm friendly smile. She pushes a vintage wooden cart with mosaic-tile wheels in blue, white, yellow, and green. The cart is filled with clearly visible everyday products with small clean English labels. Include a Talavera ceramic mug labeled "Talavera Ceramic Mug — Heat-resistant, underglaze art, 350ml"; a handwoven sisal basket labeled "Handwoven Basket — Natural sisal, sturdy, 30x20x15cm"; a green cactus-shaped scented candle with a soft flame labeled "Cactus Candle — Plant essential oil, calming, 40-hour burn"; and colorful Mexican throw pillows labeled "Mexican Pillow — Cotton-linen, removable cover, 45x45cm." The illustration background is a sunny desert oasis with real cacti, soft clouds, warm side-backlight, and a thin golden rim light on the illustrated elements. Center-right: product category navigation in a 2x2 grid of rounded rectangle cards with even spacing, 15px rounded corners, and subtle paper texture. Card 1, red background, title "Home Textiles" in bold white, with a flat icon of a woven wall tapestry with tassels; body text: "Rugs, pillows, tablecloths, curtains. Premium cotton-linen with vivid colors that bring Mexican warmth home." Card 2, peacock-blue background, title "Kitchen & Dining" in bold white, with a Talavera ceramic plate icon; body text in bright yellow: "Bowls, cups, tableware, storage. Handcrafted ceramic texture makes every meal feel special." Card 3, bright yellow background, title "Bath & Care" in bold dark brown, with a cactus-shaped soap dish icon; body text in dark brown: "Towels, soap dishes, toothbrush cups, shelves. Water-resistant materials for a fresh, relaxing bath space." Card 4, emerald-green background, title "Lifestyle Goods" in bold white, with a colorful pinwheel icon; body text in white: "Hangers, hooks, tissue boxes, trash bins. Practical, beautiful details that elevate daily living." Right side: promotional zone shaped like a large Mexican mariachi guitar. The guitar body is warm wood with red-and-gold floral vine carvings, the neck extends upward, and the strings are silvery with a metallic shine. Inside the guitar body, place the main promotion. At the top, add a red ribbon banner with white bold text: "TodoTi Grand Opening Fiesta." Below it, display a large offer: "Spend $199, Save $50," with "Spend" and "Save" in white medium text and the numbers in oversized yellow-and-red 3D type. Below the offer, add smaller handwritten white text: "New customers: $20 off first order + free shipping." At the bottom of the guitar body, place three small circular badges with dark blue backgrounds, white text, and thin gold outlines: "7-Day Easy Returns," "Fast Shipping," and "Authentic Quality." From the guitar neck, let several colorful ribbons flow outward with tilted text: "Now through the end of the month." Bottom area: brand story and culture zone against a deep blue night sky with glowing white five-pointed stars and black cactus silhouettes. On the left, place a title block: "About TodoTi" in bold white with a soft yellow glow. To the right, create three equal columns. Column 1 heading in bright yellow with a white underline: "Design Philosophy"; body in light gray-white: "TodoTi brings the color and craft of Mexican folk art into modern home essentials, blending Aztec motifs and Talavera patterns with everyday practicality." Column 2 heading: "Quality Promise"; body: "We carefully select trusted global suppliers and test every product for safe, durable, and eco-friendly quality." Column 3 heading: "Lifestyle Philosophy"; body: "Life is better in color. TodoTi helps you celebrate daily moments with warmth, joy, and vibrant Mexican-inspired energy." Very bottom: a slim interactive guidance bar above the Aztec border. Use a translucent frosted-white background with a soft blur. Left side: "Get more home inspiration from @TodoTiOfficial" in dark brown, aligned with a small official badge icon. Center: a white rounded search box with a magnifying-glass icon and placeholder text: "Search everyday essentials and start a colorful life" in warm red. Right side: a headset customer-service icon with dark brown text: "Questions? Customer Care: 400-888-TODO | 9:00–22:00." Overall style: vibrant, festive, handcrafted, authentic Mexican-inspired e-commerce banner, modern retail marketing layout, balanced composition, clean information hierarchy, high readability, warm lighting, rich details, professional storefront hero image. All on-image text must be in English only, with no Chinese characters, no misspellings, and clean legible typography.

Long paragraph: Vertical Western European 19th-century illustrated poetry page, Romantic maritime farewell theme, international English visual tradition, not an East Asian painting. The composition is a vertical art print or vintage book page with aged ivory paper, fine engraved linework, muted blue-gray watercolor washes, and elegant negative space. The main feature is a long English poem rendered clearly in elegant English copperplate-inspired calligraphy, highly legible, English alphabet only, standard left-to-right reading order, top-to-bottom layout, correct punctuation, exact line breaks, natural ink pressure, dark sepia-black lettering, subtle vintage letterpress texture, refined international typography. Below and around the text, illustrate a cold twilight seashore: gray stones beside the sea, a small boat floating near the shore, distant ships moving toward a haven under low hills, pale evening mist over the water, faint moonlight, soft ripples, a lonely coastal atmosphere, melancholic but restrained. The mood is poetic, timeless, and elegant, like a Victorian English maritime lyric. No Chinese characters, no Chinese painting style, no seals, no stamps, no modern objects, no illegible glyphs, no random letters. Render exactly the following 12 lines from Alfred, Lord Tennyson's public-domain poem "Break, Break, Break"; do not render the title or author name: Break, break, break, On thy cold gray stones, O Sea! And I would that my tongue could utter The thoughts that arise in me. O well for the fisherman's boy, That he shouts with his sister at play! O well for the sailor lad, That he sings in his boat on the bay! And the stately ships go on To their haven under the hill; But O for the touch of a vanish'd hand, And the sound of a voice that is still. High resolution, masterful long-form English text rendering, poetic, elegant, museum-quality Western illustrated book page.

Realistic portrait: A lifestyle-style male food portrait with a cinematic, on-location feel. The subject is a young man in his early twenties with clean, fresh features, gentle eyes and brows, and naturally clear skin that retains realistic texture detail. He wears a soft, off-white chunky knit sweater over a plain white T-shirt peeking out at the collar, giving off a warm, refined, and relaxed lifestyle vibe. He sits at a cozy Western restaurant table, leaning slightly forward, cutting into a Beef Wellington on his plate with a silver knife and fork in his right hand while his left hand rests naturally on the edge of the table. He tilts his head slightly toward the camera with a faint, relaxed smile and a warm, life-filled gaze. The wooden tabletop is set with an elegant Western meal, alongside a vintage stemmed glass holding an amber-colored drink and a small, delicately arranged vintage flower vase. Behind him is a clear glass wall printed with a light gold, handwritten French menu that reads "Menu du Jour," the text applied to the glass in a silkscreen-like finish with a subtle ambient reflection. The scene is lit with warm yellow overhead interior lighting that softly bathes his face and the food on the table, giving the food an appetizing sheen, while cool ambient light filtering through the glass from outside creates a cinematic contrast between warm and cool tones. The shot is a close, eye-level composition with a slight off-center angle for a natural, breathing feel rather than a stiff, centered framing. A shallow depth of field softly blurs the glass wall and restaurant setting in the background into warm bokeh, with the foreground flowers and drink edges also softly out of focus. The overall tone is warm and comforting, full of everyday atmosphere and mouthwatering appeal.

UI design: A widescreen desktop UI design for an immersive website exploring the sounds of an Asian rainforest. The overall mood is dark and tranquil, inspired by vintage field-recording equipment and analog scientific instruments. The color palette centers on deep emerald green, rich obsidian black, and soft earthy brown, creating a calm and serene atmosphere. The interface features subtle mist and haze effects, with soft, dappled light refraction simulating sunlight filtering through a dense canopy, and the background incorporates natural textures such as fine paper grain and leaf veins. The overall image is highly realistic, in a style comparable to a premium BBC nature documentary, with a cinematic composition and extremely fine rendering detail. A minimal, semi-transparent navigation bar spans the very top of the page. On the far left is the brand mark, formed by an intertwined line-art illustration of a vintage analog microphone and monstera leaves, paired with the brand name "RAINFOREST ARCHIVE" in an elegant, soft cream-colored serif font. On the right side of the navigation bar is a menu with four minimal text links in a matching font and color: "Expeditions," "Species," "Field Notes," and "About." The upper two-thirds of the page is the main visual hero banner. The background is an ultra-high-definition, richly atmospheric photo of an ancient Southeast Asian dipterocarp rainforest, wrapped in soft, drifting mist, with blurred dark-green fern fronds in the foreground creating a natural vignette. Overlaid on the left of the banner is the main headline, in a large, high-contrast, elegant serif font: "Voices of the Canopy." Directly below the headline is a smaller, refined subtitle in light cream: "An auditory journey through the ancient dipterocarp forests of Southeast Asia." The text has a subtle typewriter texture, with slightly irregular character spacing that reinforces the vintage, analog-equipment feel. In the lower-middle of the banner is the core audio interaction component: a precision player modeled after a vintage field-recording reel-to-reel deck, inspired by classic Nagra or Marantz tape recorders. The player body is brushed dark gunmetal-gray metal with subtle natural scratches and wear marks, and features two large knurled metal knobs printed with small, clear sans-serif labels: "GAIN" and "MONITOR." At the center of the device is a classic VU meter with a warm, soft amber backlight, its needle resting near the zero mark. To the right of the meter is a heavy-duty physical toggle switch, flipped up to indicate "PLAY" status. While audio is playing, the VU meter emits slowly expanding concentric ripples, resembling gentle water ripples or faint sound waves, in a translucent misty green tone that blends naturally with the hazy background. Below the banner, the page layout shifts to a clean, documentary-style grid used as the sound library section. At the top of this section is an elegant small heading in wide-spaced, all-caps type: "CURRENT EXPEDITIONS." Below the heading is a tidy two-column grid displaying four soundscape cards, each with a distinct style. Card 1 (top left): background is a moody, misty riverbank thumbnail, with the title "Dawn Chorus in Danum Valley," and below it the caption "Location: Sabah, Borneo" and "Duration: 42:15." A small amber badge on the right of the card reads "Currently Playing." Card 2 (top right): a close-up of wet, broad green leaves, with the title "Monsoon Rain on Broadleaves," caption "Location: Khao Yai, Thailand" and "Duration: 1:15:30," and a thin-outlined button labeled "Listen." Card 3 (bottom left): a hazy silhouette of the forest canopy at dusk, with the title "Nocturnal Gibbon Calls," caption "Location: Khao Sok, Thailand" and "Duration: 58:02," and a thin-outlined button labeled "Listen." Card 4 (bottom right): a macro close-up of tree bark texture, with the title "Hornbill Wingbeats," caption "Location: Taman Negara, Malaysia" and "Duration: 24:40," and a thin-outlined button labeled "Listen." On the far right of the soundscape grid, a vertical panel spans the lower half of the section, styled after a page from a vintage field science notebook. The paper has a soft, aged cream texture, with a faint coffee stain in the bottom-right corner. The panel title uses a typewriter monospace font reading "FIELD NOTES," with elegant, slightly faded ink-toned body text: "Recorded using parabolic microphones and Nagra IV-S tape recorders. Humidity: 94% Temp: 24°C." Next to the text is a finely hand-drawn dark brown ink sketch of a pitcher plant. At the very bottom of the page is a minimal footer spanning the full width, with a thin dark green divider line above it. The footer reads "Copyright 2024 Rainforest Archive" on the left and "Funded by the Wildlife Conservation Society" on the right, both in a soft, understated small serif font.

PPT: A professional educational PPT slide with a 16:9 aspect ratio. The background is an extremely subtle blue-gray gradient, overlaid with a very low-opacity sine-wave pattern and a faint dot-grid texture, creating a rigorous academic physics atmosphere with a tech-forward feel. Centered at the top of the slide is the main title area, with the text in large, bold, dark navy sans-serif type: "AC Circuit Analysis: Resistor with Sinusoidal Voltage." Directly below the title is a bright thin orange horizontal line as a visual divider, with small circular end caps at both ends to add a sense of ceremony to the layout. The main body of the slide is split into two columns using a strict grid-alignment system, with the left and right columns in roughly a 1:1.2 width ratio, giving the complex formulas on the right ample display space. At the top of the left column is a rounded-rectangle orange header bar with centered bold white text: "PROBLEM DATA & WAVEFORM." The header bar has a subtle drop shadow at the bottom, giving it a slightly raised, three-dimensional feel. Below the orange header is a light gray-white rounded data card with an extremely thin light blue border, containing four vertically arranged data points. Each point has a minimalist, solid-gray line icon to its left — in order, a resistor symbol, a wave line, a voltage arrow, and a clock symbol. The text for the four points is in clear dark gray type, reading in order: "Resistor (R) = 10 Ω," "Waveform = Sinusoidal," "Peak Voltage (V subscript p) = 311 V," and "Period (T) = 0.02 s," neatly laid out with comfortable line spacing, with key numbers highlighted in a slightly darker color. Below the data card is the chart area, with a dark gray chart title centered at the top: "Voltage vs. Time." The chart itself is a 2D coordinate system with light gray gridlines. To the left of the vertical Y-axis is the dark gray label "Voltage (V)," with tick marks labeled "+311," "0," and "-311" in dark blue text on white background boxes. Below the horizontal X-axis is the dark gray label "Time (s)," with tick marks labeled "0.01" and "0.02." At the center of the coordinate system is a smooth, full, slightly glowing bright red sine-wave curve of moderate line weight, showing perfect periodicity. At the peak of the sine wave, a dark blue vertical double-headed arrow indicates the amplitude, labeled in dark blue text: "V subscript p = 311 V." Above one full period on the X-axis is a dark blue horizontal dimension line with vertical lead lines at both ends, labeled in the center in dark blue text: "T = 0.02 s." The chart area has generous white space overall, keeping the visual focus on the red waveform. At the top of the right column is another rounded-rectangle orange header bar, aligned in height with the left header, with centered bold white text: "SOLUTION STEPS & CALCULATIONS." Below the orange header, four light-blue gradient rounded cards are arranged vertically with even spacing between them and generous internal breathing room. The first card has a dark blue circular numbered badge in the top-left corner with a white "1" inside. The card title is bold dark blue text: "1. Maximum Current (I subscript p)." The body below is laid out in three lines in the visual style of standard mathematical notation — with fraction bars, sub/superscripts, and radicals rendered elegantly — reading in dark gray: "Formula: Ohm's Law, I subscript p = V subscript p over R," followed by the calculation "I subscript p = 311 V over 10 Ω = 31.1 A." In the bottom-right corner of the card is a bright yellow rounded highlight box with bold dark blue text: "I subscript p = 31.1 A." The second card's badge shows "2," titled "2. AC Voltmeter Reading (V subscript rms)." The body includes the note "Voltmeters read RMS value," the formula "V subscript rms = V subscript p over square root of 2," and the calculation "V subscript rms = 311 V over 1.414 ≈ 220 V," with the bottom-right highlight box showing "V subscript rms ≈ 220 V." The third card's badge shows "3," titled "3. Actual Power Dissipated (P subscript avg)." The body includes "Using RMS values," the formula "P subscript avg = (V subscript rms) squared over R," the calculation "P subscript avg = (220 V) squared over 10 Ω = 48400 over 10 W = 4840 W," with the bottom-right highlight box showing "P subscript avg = 4.84 kW." The fourth card's badge shows "4," titled "4. Joule Heat in Half Cycle (Q subscript half)." The body includes "Energy = Power × Time," the time calculation "Time = T over 2 = 0.01 s," the main calculation "Q subscript half = P subscript avg × (0.01 s) = 4840 W × 0.01 s = 48.4 J," with the bottom-right highlight box showing "Q subscript half = 48.4 J." All four cards' highlight boxes are aligned uniformly in the bottom-right corner, creating a tidy visual alignment and reading flow. At the very bottom of the slide is an extremely thin light gray horizontal divider, below which is centered small light gray sans-serif text: "Experimental Physics - AC Circuits Analysis Module." The footer area remains minimal, with the rest of the background kept as clean graphics and white space. The slide's overall color system is built around dark navy, bright orange, light blue-gray, and pure white, with the red waveform and yellow highlight boxes as visual accents that break the monotony and add vibrancy and professionalism. All text has sharp edges and extremely high contrast, ensuring clear readability on projectors or screens. Lighting and shadow are handled with restraint, adding depth only through subtle drop shadows beneath the cards and slight highlights on the header bars, resulting in an overall highly polished, clearly structured, and information-rich professional physics teaching slide.

Illustration design: A hand-drawn watercolor illustration on soft, off-white watercolor paper. The paper surface shows the characteristics of traditional botanical art, including clearly visible pigment granulation and the soft color bleeding produced by wet-on-wet technique. In the upper-center area of the composition is a cluster of violets and pansies. The petals blend gradients of light purple, deep purple, and lavender, with golden-yellow flower centers. Beside the flowers are small, emerald-green leaves rendered with translucent watercolor washes and finely detailed veins. Soft directional lighting illuminates the scene, highlighting the texture of the watercolor paper and the subtle variations in the pigment. Directly below the floral motif, horizontally centered in the lower half of the paper, is the word "Виолета." The lettering is written in an arched calligraphic script using metallic gold ink. The letterforms are rendered with fluid, continuous strokes and a subtle embossed, three-dimensional quality, giving off a soft metallic sheen under the light.

Wanxiang

Portrait photography

p1023408

Photorealistic photography

p1023409

Artistic styles

p1023411

Text rendering

p1023399

Poster design

image.png

Image set generation

p1023424

View prompt

Portrait photography: A realistic portrait photograph. Background: the red walls of the Forbidden City. A woman in a black cheongsam holds a fan. Use long-exposure photography to create a cinematic, story-driven feel inspired by Wong Kar-wai. Passersby blur into motion trails, and dreamy lighting creates soft, flowing light paths. Use a soft-focus effect. The woman's gaze is deep and enigmatic, conveying a strong artistic presence.

Photorealistic photography: Photorealistic photography of a fox staring into the lens in a forest. A fisheye perspective creates strong distortion. Fur details are sharp, and background trees are warped into a circular pattern. Watercolor style with soft tones.

Artistic styles: A bouquet of wildflowers in an old earthenware vase. Background: a country kitchen. Impressionist style, soft brushstrokes, warm light, oil painting texture.

Text rendering: Traditional Chinese ink brush painting with visible Xuan paper texture. A light ink wash outlines a hazy living room. An Eastern girl in a plain dress sits cross-legged on a softly blurred vintage fabric sofa, her head lowered as she holds an unrolled poem scroll. Outside, bamboo shadows sway and a light breeze stirs the curtains. The composition has generous white space. On the right, small regular script reads: "Sitting idle, grieving at graying temples; dreams drift into blue smoke." A vermilion seal is in the lower-left corner. Ink tones vary naturally, with dry-brush strokes creating a sense of flowing light and shadow. The mood is serene and profound, evoking the lingering notes of a guqin.

Poster Design: Flat geometric illustration style, a Dragon Boat Festival poster, magazine cover. Color Palette and Background: The main color is a pink gradient, creating a soft and festive background that sets a warm and traditional tone. Text Elements: Green font with a shadow effect. The main copy highlights "DRAGONBOAT FESTIVAL" and "Duanwu" on two separate lines. Below the body text, "2025/05/31" and "the fifth day of the fifth lunar month" highlight the date "2025/05/31". Main Graphic: A dragon boat with a green body and pink fins, in highly saturated colors with strong contrast, surrounded by auspicious cloud elements. There are figures on the boat to further evoke the scene of dragon boat racing and add festive energy. Detail Embellishments: Add the text "Chinese Traditional Festival" paired with small zongzi icons to enrich the cultural details. Advanced minimalist layout, a masterpiece. Simple, stylish, and grand, a new Chinese-style traditional poster. The font should not have a shadow style.

Image set generation: A four-panel, Japanese-style chibi manga in a cel-shading style. Panel 1: A programmer with black-rimmed glasses stares at a red error message on the screen. His pupils shrink in shock as cold sweat drips down. The background cracks and dissolves into a pixelated abyss. Panel 2: He rolls up his sleeves, types confidently, and raises an eyebrow. A speech bubble appears above him: "This is just a five-line code fix!" Panel 3: His screen is flooded with chaotic error symbols. His hair stands on end, dark circles ring his eyes, and his chair tilts back 45 degrees. Discarded flowcharts float across the ceiling. Panel 4: After accidentally deleting a single grayed-out comment line, a green checkmark flashes. He tilts his head in a daze as a question bubble appears on the screen: "...So it was all just a hallucination?"

Model selection

  • qwen-image-3.0-pro: The flagship Qwen Image 3.0 model. Supports intelligent prompt rewriting and excels at rendering text that blends naturally with physical materials in both Chinese and English.

  • wan2.7-image-pro: Offers the most features, including multi-image generation, resolutions up to 4096x4096, and enhanced control over facial features, colors, and long text rendering.

  • z-image-turbo: Delivers fast, cost-effective image generation, excelling at highly realistic portraits and product images.

Quick start

Prerequisites

Before you begin, get an API key, then set the API key as an environment variable. If you use the DashScope SDK, you must also install the SDK.

Sample code

Calling methods:

Qwen - synchronous call

Python

import os
import dashscope
from dashscope import MultiModalConversation

dashscope.base_http_api_url = 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1'

response = MultiModalConversation.call(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    model="qwen-image-3.0-pro",
    messages=[{
        "role": "user",
        "content": [
            {"text": "A vertical outdoor portrait photograph with a warm afternoon street atmosphere. Deep green vines and small orange flowers cascade from building eaves across the upper area. A dark blue sign reads 'Il Messaggero' in white Gothic lettering, partially obscured by foliage. Below, a newsstand displays newspapers behind black metal-framed glass, blurred by shallow depth of field. Strong backlight streams from the street's end. Center-right, a young woman in a black spaghetti-strap backless dress looks back at the camera with a warm smile. Her long, thick wavy black hair is outlined by golden rim light. She has fair skin, bright eyes, soft coral-red lips, and holds a large bouquet of orange, apricot, pink and peach roses contrasting with her black dress. The sunlit city street stretches into the blurred background. Warm film-like tones with fine grain, soft contrast and pronounced backlit edge glow create a romantic, bright, urban strolling atmosphere."}
        ]
    }],
    prompt_extend=True
)

print(response)
if response.status_code == 200:
    url = response.output.choices[0].message.content[0]["image"]
    print(f"Generated image URL: {url}")
else:
    print(f"Error: {response.code} - {response.message}")

Java

import java.util.Arrays;
import java.util.Collections;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.utils.Constants;

public class ImageEditExample {
    public static void main(String[] args) {
        Constants.baseHttpApiUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";

        MultiModalConversation conv = new MultiModalConversation();
        MultiModalMessage userMessage = MultiModalMessage.builder()
            .role(Role.USER.getValue())
            .content(Arrays.asList(
                Collections.singletonMap("text", "A vertical outdoor portrait photograph with a warm afternoon street atmosphere. Deep green vines and small orange flowers cascade from building eaves across the upper area. A dark blue sign reads 'Il Messaggero' in white Gothic lettering, partially obscured by foliage. Below, a newsstand displays newspapers behind black metal-framed glass, blurred by shallow depth of field. Strong backlight streams from the street's end. Center-right, a young woman in a black spaghetti-strap backless dress looks back at the camera with a warm smile. Her long, thick wavy black hair is outlined by golden rim light. She has fair skin, bright eyes, soft coral-red lips, and holds a large bouquet of orange, apricot, pink and peach roses contrasting with her black dress. The sunlit city street stretches into the blurred background. Warm film-like tones with fine grain, soft contrast and pronounced backlit edge glow create a romantic, bright, urban strolling atmosphere.")
            ))
            .build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
            .apiKey(System.getenv("DASHSCOPE_API_KEY"))
            .model("qwen-image-3.0-pro")
            .messages(Arrays.asList(userMessage))
            .parameter("prompt_extend", true)
            .build();
        try {
            MultiModalConversationResult result = conv.call(param);
            System.out.println(result);
        } catch (Exception e) {
            e.printStackTrace();
        }
    }
}

curl

Request example
curl --location 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header 'Content-Type: application/json' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--data '{
    "model": "qwen-image-3.0-pro",
    "input": {
        "messages": [
            {
                "role": "user",
                "content": [
                    {
                        "text": "A vertical outdoor portrait photograph with a warm afternoon street atmosphere. Deep green vines and small orange flowers cascade from building eaves across the upper area. A dark blue sign reads '\''Il Messaggero'\'' in white Gothic lettering, partially obscured by foliage. Below, a newsstand displays newspapers behind black metal-framed glass, blurred by shallow depth of field. Strong backlight streams from the street'\''s end. Center-right, a young woman in a black spaghetti-strap backless dress looks back at the camera with a warm smile. Her long, thick wavy black hair is outlined by golden rim light. She has fair skin, bright eyes, soft coral-red lips, and holds a large bouquet of orange, apricot, pink and peach roses contrasting with her black dress. The sunlit city street stretches into the blurred background. Warm film-like tones with fine grain, soft contrast and pronounced backlit edge glow create a romantic, bright, urban strolling atmosphere."
                    }
                ]
            }
        ]
    },
    "parameters": {
        "prompt_extend": true
    }
}'
Response example
{
    "output": {
        "choices": [
            {
                "finish_reason": "stop",
                "message": {
                    "content": [
                        {
                            "image": "https://dashscope-result-sz.oss-cn-shenzhen.aliyuncs.com/xxx.png?Expires=xxx"
                        }
                    ],
                    "role": "assistant"
                }
            }
        ]
    },
    "usage": {
        "output_height": 1024,
        "output_width": 1024,
        "input_image_count": 1,
        "input_image_type": "qima_input_1k",
        "output_image_count": 1,
        "output_image_type": "qima_output_1k"
    },
    "request_id": "571ae02f-5c9d-436c-83c2-f221e6df0xxx"
}

Wan - asynchronous call

Python

Request example
import os
import dashscope
from dashscope.aigc.image_generation import ImageGeneration
from dashscope.api_entities.dashscope_response import Message

# This is the base URL for the China (Beijing) region; base URLs are region-specific.
dashscope.base_http_api_url = 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1'

# If the environment variable is not set, replace the following line with your Model Studio API key: api_key="sk-xxx"
# API keys vary by region. To get an API key, visit: https://help.aliyun.com/en/model-studio/get-api-key
api_key = os.getenv("DASHSCOPE_API_KEY")


def main():
    message = Message(
        role="user",
        content=[
            {"text": "A young woman in a natural, casual selfie style. An ultra-high-definition, realistic lifestyle photo. She is wearing a yellow floral long-sleeved top, and her long, slightly wavy hair falls naturally. The background is an outdoor natural scene with green plants nearby and water and mountains in the distance. Soft, natural sunlight falls on her face and body, creating natural light and shadow effects. The camera angle is a medium shot from a selfie perspective, as if held by her. She is standing naturally, projecting a relaxed and comfortable state. The angle is natural, in the style of a casual snapshot—an unguarded moment."}
        ]
    )
    
    # Submit an asynchronous task.
    print("Submitting the asynchronous task...")
    response = ImageGeneration.async_call(
        model="wan2.7-image-pro",
        api_key=api_key,
        messages=[message],
        enable_sequential=False,
        n=1,
        size="2K"
    )
    
    if response.status_code == 200:
        print(f"Task submitted successfully. Task ID: {response.output.task_id}")
        
        # Wait for the task to complete.
        status = ImageGeneration.wait(task=response, api_key=api_key)
        
        if status.output.task_status == "SUCCEEDED":
            print("Task completed!")
            print(f"Result:")
            print(status)
        else:
            print(f"Task failed. Status: {status.output.task_status}")
    else:
        print(f"Task creation failed: {response.code} - {response.message}")


if __name__ == "__main__":
    try:
        main()
    except Exception as e:
        print(f"Error: {e}")
Response example

1. Task creation response

{
    "status_code": 200,
    "request_id": "4fb3050f-de57-4a24-84ff-e37ee5xxxxxx",
    "code": "",
    "message": "",
    "output": {
        "text": null,
        "finish_reason": null,
        "choices": null,
        "audio": null,
        "task_id": "77093787-a217-4c29-9cd4-ca7b5ac86xxx",
        "task_status": "PENDING"
    },
    "usage": {
        "input_tokens": 0,
        "output_tokens": 0,
        "characters": 0
    }
}

2. Task status query response

Note

The usage field is included in the raw HTTP API response. In the Python SDK, the ImageGeneration.wait() method returns a response object where the output property contains task status fields (such as task_status, task_id, choices, submit_time, scheduled_time, end_time, finished), but does not include the usage field. To access billing information (such as image_count, total_tokens, size), use the raw HTTP response or query the task status API directly.

The image URL is valid for 24 hours. Download the image promptly.
{
    "status_code": 200,
    "request_id": "56e318fd-ed60-99e8-8ca1-cdef25ca4xxx",
    "code": "",
    "message": "",
    "output": {
        "text": null,
        "finish_reason": null,
        "choices": [
            {
                "finish_reason": "stop",
                "message": {
                    "role": "assistant",
                    "content": [
                        {
                            "image": "https://dashscope-result-bj.oss-cn-beijing.aliyuncs.com/xxxxxx.png?Expires=xxxxxx",
                            "type": "image"
                        }
                    ]
                }
            }
        ],
        "audio": null,
        "task_id": "77093787-a217-4c29-9cd4-ca7b5ac86xxx",
        "task_status": "SUCCEEDED",
        "submit_time": "2026-03-31 23:04:46.166",
        "scheduled_time": "2026-03-31 23:04:46.208",
        "end_time": "2026-03-31 23:05:11.664",
        "finished": true
    },
    "usage": {
        "input_tokens": 720,
        "output_tokens": 11,
        "characters": 0,
        "size": "2048*2048",
        "total_tokens": 731,
        "image_count": 1
    }
}

Java

Request example
import com.alibaba.dashscope.aigc.imagegeneration.*;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.JsonUtils;

import java.util.Collections;

public class Main {

    // This is the base URL for the China (Beijing) region; base URLs are region-specific.
    static String baseUrl = "https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1";

    // If the environment variable is not set, replace the following line with your Model Studio API key: apiKey="sk-xxx"
    // API keys vary by region. To get an API key, visit: https://help.aliyun.com/en/model-studio/get-api-key
    static String apiKey = System.getenv("DASHSCOPE_API_KEY");

    public static ImageGenerationResult waitTask(String taskId)
            throws ApiException, NoApiKeyException {
        ImageGeneration imageGeneration = new ImageGeneration(apiKey, baseUrl);
        return imageGeneration.wait(taskId, apiKey);
    }

    public static void asyncCall() throws ApiException, NoApiKeyException, UploadFileException {
        ImageGenerationMessage message = ImageGenerationMessage.builder()
                .role("user")
                .content(Collections.singletonList(
                        Collections.singletonMap("text", "A young woman in a natural, casual selfie style. An ultra-high-definition, realistic lifestyle photo. She is wearing a yellow floral long-sleeved top, and her long, slightly wavy hair falls naturally. The background is an outdoor natural scene with green plants nearby and water and mountains in the distance. Soft, natural sunlight falls on her face and body, creating natural light and shadow effects. The camera angle is a medium shot from a selfie perspective, as if held by her. She is standing naturally, projecting a relaxed and comfortable state. The angle is natural, in the style of a casual snapshot—an unguarded moment.")
                )).build();

        ImageGenerationParam param = ImageGenerationParam.builder()
                .apiKey(apiKey)
                .model("wan2.7-image-pro")
                .messages(Collections.singletonList(message))
                .enableSequential(false)
                .n(1)
                .size("2K")
                .build();

        ImageGeneration imageGeneration = new ImageGeneration(apiKey, baseUrl);
        ImageGenerationResult taskResult = null;
        try {
            System.out.println("----async call, creating task----");
            taskResult = imageGeneration.asyncCall(param);
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            throw new RuntimeException(e.getMessage());
        }
        System.out.println("Task created: " + JsonUtils.toJson(taskResult));

        // Wait for the task to complete.
        String taskId = taskResult.getOutput().getTaskId();
        ImageGenerationResult result = waitTask(taskId);
        System.out.println(JsonUtils.toJson(result));
    }

    public static void main(String[] args) {
        try {
            asyncCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
    }
}
Response example

1. Task creation response

{
    "requestId": "7d026dc1-e8c9-9caa-84ac-e82e2da97xxx",
    "output": {
        "task_id": "2de18c56-c151-4b80-8105-1d164733exxx",
        "task_status": "PENDING"
    },
    "status_code": 200,
    "code": "",
    "message": ""
}

2. Task status query response

{
    "requestId": "daea7295-4ce0-928a-9a11-4d2bea058xxx",
    "usage": {
        "input_tokens": 720,
        "output_tokens": 11,
        "total_tokens": 731,
        "image_count": 1,
        "size": "2048*2048"
    },
    "output": {
        "choices": [
            {
                "finish_reason": "stop",
                "message": {
                    "role": "assistant",
                    "content": [
                        {
                            "image": "https://dashscope-result-bj.oss-cn-beijing.aliyuncs.com/xxxxxx.png?Expires=xxxxxx",
                            "type": "image"
                        }
                    ]
                }
            }
        ],
        "task_id": "2de18c56-c151-4b80-8105-1d164733exxx",
        "task_status": "SUCCEEDED",
        "finished": true,
        "submit_time": "2026-03-31 19:49:53.124",
        "scheduled_time": "2026-03-31 19:49:53.175",
        "end_time": "2026-03-31 19:50:53.160"
    },
    "status_code": 200,
    "code": "",
    "message": ""
}

Curl

Note
  • For an asynchronous call, you must set the Header parameter X-DashScope-Async to enable.

  • The task_id of an asynchronous task can be queried for 24 hours. After this period, the task status changes to UNKNOWN.

  • This method works for all models. For beginners, we recommend using Postman to call the API.

Step 1: Create a task

The request returns a task ID (task_id).

curl --location 'https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/services/aigc/image-generation/generation' \
--header 'Content-Type: application/json' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header "X-DashScope-Async: enable" \
--data '{
    "model": "wan2.7-image-pro",
    "input": {
        "messages": [
            {
                "role": "user",
                "content": [
                    {"text": "A flower shop with exquisite windows, a beautiful wooden door, and flowers on display"}
                ]
            }
        ]
    },
    "parameters": {
        "size": "2K",
        "n": 1,
        "watermark": false,
        "thinking_mode": true
    }
}'
Step 2: Query task result

Use the task_id from the previous step to poll the API for the task status until the task_status changes to SUCCEEDED or FAILED.

Replace {task_id} with the task_id value returned by the previous API call. The task_id is valid for queries for 24 hours, Replace {WorkspaceId} with your actual workspace ID.

curl -X GET https://{WorkspaceId}.cn-beijing.maas.aliyuncs.com/api/v1/tasks/{task_id} \
--header "Authorization: Bearer $DASHSCOPE_API_KEY"

Key capabilities

1. Prompt following

Parameters: messages.content.text or input.prompt (required), and negative_prompt (optional).

  • prompt: Describes the desired content for the image, including the subject, scene, style, lighting, and composition. This is the core parameter for controlling text-to-image generation.

  • negative prompt: Describes what to exclude from the image, such as "blurry" or "extra fingers." This parameter helps refine the output quality.

For best results, use a structured prompt. For more information, see Text-to-image prompt guide.

The negative_prompt parameter is not supported by wan2.7-image-pro and wan2.7-image. To exclude unwanted elements, describe them in the prompt (for example, "do not include xxx").

2. Enable prompt rewriting

Parameter: parameters.prompt_extend (Boolean, defaults to true).

This feature automatically expands and optimizes short prompts to improve image quality. Enabling this feature adds 3 to 5 seconds to the generation time, as a large model is used to rewrite the prompt.

Recommendations:

  • Enable this feature for brief or general prompts to significantly improve image quality.

  • Disable this feature if you need precise control over image details, have provided a comprehensive prompt, or if response latency is a concern. To disable it, set the prompt_extend parameter to false.

The prompt_extend parameter is not supported by wan2.7-image-pro and wan2.7-image. To improve image quality for these models, enable thinking_mode instead. The qwen-image-3.0-pro and qwen-image-3.0 models support the prompt_extend parameter (enabled by default).

Parameter: parameters.prompt_extend_mode (string, optional, defaults to direct). Supported only by qwen-image-3.0-pro and qwen-image-3.0. Specifies the rewriting method used when prompt_extend is enabled:

  • direct: Direct Prompt Enhancement (DPE), suitable for most scenarios. Supported for both text-to-image (T2I) and image-to-image/editing (I2I).

  • agent: Agent Prompt Enhancement (APE), provides more refined rewriting. Only supported for text-to-image (T2I). Passing agent for image-to-image/editing (I2I) returns a 400 error.

3. Set the output image resolution

Parameter: parameters.size (string), in the "width*height" format.

Model

Size format

Pixel range

Default

Aspect ratio

wan2.7-image-pro

Shorthand

1K (1024*1024), 2K (2048*2048), 4K (4096*4096)

2K (2048*2048)

1:8 – 8:1

Custom "width*height"

768*768 – 4096*4096

wan2.7-image

Shorthand

1K (1024*1024), 2K (2048*2048)

2K (2048*2048)

1:8 – 8:1

Custom "width*height"

768*768 – 2048*2048

wan2.6-image (interleaved text-image output mode)

Custom "width*height"

768*768 – 1280*1280

Matches input aspect ratio (≤1280*1280)

1:4 – 4:1

wan2.6-t2i, wan2.5-t2i-preview

Custom "width*height"

1280*1280 – 1440*1440

1280*1280

1:4 – 4:1

wan2.2 and earlier t2i models

Custom "width*height"

[512, 1440], and total pixels ≤1440*1440

1024*1024

-

qwen-image-3.0-pro, qwen-image-3.0

Custom "width*height"

512*512 – 2048*2048

Auto-recommended by model

1:8 – 8:1

qwen-image-2.0 series

Custom "width*height"

512*512 – 2048*2048

2048*2048

-

qwen-image-max / qwen-image-plus series

Fixed preset sizes only

See preset sizes below

1664*928 (16:9)

-

The wan2.7-image-pro model supports 4K resolution and custom resolutions up to 4096*4096, but only for text-to-image tasks (where no image is input and image set generation is disabled). All other scenarios are limited to 2K resolution (2048*2048).

The qwen-image-max and qwen-image-plus series support only the following five fixed resolutions:

  • 1664*928 (default): 16:9

  • 1472*1104: 4:3

  • 1328*1328: 1:1

  • 1104*1472: 3:4

  • 928*1664: 9:16

Recommended resolutions:

Aspect ratio

4K (wan2.7-image-pro)

2K (wan2.7-image, qwen-image-3.0-pro, qwen-image-3.0, qwen-image-2.0)

1K (Wan t2i)

1:1

4096*4096

2048*2048

1280*1280

16:9

4096*2304

2688*1536

1696*960

9:16

2304*4096

1536*2688

960*1696

4:3

4096*3072

2368*1728

1472*1104

3:4

3072*4096

1728*2368

1104*1472

4. Image set generation

Parameter: parameters.enable_sequential (Boolean, defaults to false). Supported only by wan2.7-image-pro and wan2.7-image.

Set to true to enable image set generation mode. In this mode, the model uses the prompt and any reference images to generate multiple, story-coherent images from a single request.

  • Number of images: Controlled by the n parameter. When this mode is enabled, this value can range from 1 to 12, with a default of 12. The model determines the actual number of images it generates, which will not exceed n.

  • Note: When image set generation is enabled, the thinking_mode and color_palette parameters are unavailable.

5. Thinking mode

Parameter: parameters.thinking_mode (Boolean, defaults to true). Supported only by wan2.7-image-pro and wan2.7-image.

When enabled, the model enhances its reasoning capabilities to improve image quality. This increases the generation time.

Available only when image set generation is disabled (enable_sequential=false).

6. Custom color palette

Parameter: parameters.color_palette (array). Supported only by wan2.7-image-pro and wan2.7-image.

Define the image's color scheme by providing an array of objects, where each object specifies a hex color and its ratio. You must provide 3 to 10 colors (8 is recommended), and the sum of all ratios must equal 100.00%.

Available only when image set generation is disabled (enable_sequential=false).

Click to view an input example

"color_palette": [
    {
        "hex": "#C2D1E6",
        "ratio": "23.51%"
    },
    {
        "hex": "#CDD8E9",
        "ratio": "20.13%"
    },
    {
        "hex": "#B5C8DB",
        "ratio": "15.88%"
    },
    {
        "hex": "#C0B5B4",
        "ratio": "13.27%"
    },
    {
        "hex": "#DAE0EC",
        "ratio": "10.11%"
    },
    {
        "hex": "#636574",
        "ratio": "8.93%"
    },
    {
        "hex": "#CACAD2",
        "ratio": "5.55%"
    },
    {
        "hex": "#CBD4E4",
        "ratio": "2.62%"
    }
]

Production deployment

  • Fault tolerance

    • Handling rate limiting: When the API returns the Throttling error code or the HTTP 429 status code, rate limiting has been triggered. To handle rate limiting, see Rate Limiting.

    • Polling for asynchronous tasks: When polling for the result of an asynchronous task, implement a reasonable polling strategy to avoid triggering rate limiting. For example, poll every 3 seconds for the first 30 seconds, then increase the polling interval. Set a final timeout for the task (e.g., 2 minutes). If the task times out, mark it as failed.

  • Risk prevention

    • Persist results: The API's image URLs are valid for 24 hours. Your production system must download the image immediately after receiving the URL and transfer it to your own persistent storage service, such as Alibaba Cloud Object Storage Service (OSS).

    • Content moderation: All prompt and negative_prompt inputs are subject to content moderation. If the input is non-compliant, the request is blocked and a DataInspectionFailed error is returned.

    • Copyright and compliance risks of generated content: Ensure that your prompts comply with all applicable laws and regulations. Generating content that includes brand trademarks, celebrity likenesses, or copyrighted intellectual property (IP) may pose infringement risks. You are responsible for evaluating and bearing all associated risks.

API reference

Billing and rate limiting

Error codes

If the model call fails and returns an error message, see Error codes for resolution.

FAQ

Q: How long are image URLs valid? How do I save images permanently?

A: Image URLs expire after 24 hours. You must programmatically download the image immediately and save it to persistent storage, such as a local server or Alibaba Cloud Object Storage Service.

Q: My API call returns a DataInspectionFailed error. How do I resolve this?

A: This error means your input triggered content moderation. Review your prompt or negative_prompt, remove any non-compliant content, and then retry the request.

Q: When should I enable or disable the prompt_extend parameter?

A: Keep it enabled (the default) for concise prompts or for more creative output. Set it explicitly to false when your prompt is already detailed and specialized, or when you are sensitive to API latency.

Note: The wan2.7-image-pro and wan2.7-image models do not support the prompt_extend parameter. To improve image quality, enable thinking_mode instead.