Premium Without the AI Look · The Journal

How to create an AI model for your e-commerce brand

Five steps to a face your brand owns, and every prompt written out and ready to paste.

By Anne Reyes · · 8 min read

Your store can have the same model in every photo. I trained her once and she never changes.

This is the file behind that video. The five steps, then every prompt I use, written out so you can paste them straight in and swap what belongs to your brand.

An AI brand model in an oversized camel blazer against a seamless white studio backdrop, looking directly into the lens
The modelOne trained identity. Every image after this one is her.
What is inside The five steps. The casting prompt. The wardrobe clause that controls what she wears. How to take a look you like and turn it into a prompt. The five-angle set that turns one image into a usable library. And the video prompt, with the settings that make it work.

What are the five steps?

  1. Open Claude.
  2. Connect the Higgsfield connector. Claude then drives the image and video models directly, so you are writing briefs rather than clicking through an interface.
  3. Upload 40 to 80 photos of the same person. Different angles, different light, same face in every frame.
  4. Tell Claude to train a Soul Character, and give her a name. Mine is called Anne-Model-01. Training runs in the background and takes about ten minutes.
  5. Generate with Soul 2.0 and call that model in every prompt from then on. The trained identity only works with Soul 2.0 and Soul Cinema. No other model will accept it.

Once training finishes you get back a soul_id. Save it somewhere you will find it again. That string is the asset. Everything else on this page is reproducible in an afternoon; that is not.

The one rule that breaks it Every reference photo has to be the same person. A mixed set does not produce a versatile model, it produces a blurry average of several people, and no amount of prompting afterwards will separate them again. If you do not have enough photos of one consistent face, generate them: run the casting prompt below, pick the best output, then generate the rest from that image.

The casting prompt

This is the one that decides who she is, so it is the one worth spending time on. Swap the age, coloring, hair and wardrobe for your customer. Leave the lighting, lens and skin-texture language alone, because that is what makes it read as photography rather than as a render.

Prompt 1 · casting
Editorial beauty portrait of a woman in her late 20s, straight-on and symmetrical, direct unsmiling gaze into the lens, chin level, composed and still. Natural skin texture with visible pores, light freckles across the nose and cheeks, dewy skin with a soft sheen, warm brown eyeshadow with a fine dark lash line, thick full natural brows, nude satin lip. Voluminous dark brown hair pushed back off the face with height and volume at the crown, loose undone waves falling behind the shoulders. Wearing an oversized camel tailored blazer in matte fine wool, worn open directly against bare skin with no shirt, top or camisole underneath, one hand raised to hold the lapel gathered closed at the chest, small brown buttons visible on the sleeve cuff, wide notch lapels, collarbone and neck exposed. Broad soft frontal light from a large diffused source, near shadowless, high key. Seamless white studio backdrop. Shot on 85mm at f2, shallow depth of field, head and chest crop. Muted warm neutral color grade, camel and cream palette. Tack sharp on the eyes.

Model Soul 2.0  ·  Size 1344 × 2016  ·  Quality 1080p  ·  Prompt enhancement off  ·  No attached reference image

Run it with four variations rather than one. You are casting, and casting means seeing options next to each other. Then pick one and be able to say why: skin texture, a face that is not perfectly symmetrical, and whether she looks like she has a job.

The wardrobe clause

Wardrobe is where most people lose control of the image, because these models fill in whatever they expect to see. An open blazer quietly acquires a shirt underneath. The fix is not a negative instruction, it is an exhaustive positive one.

Prompt 2 · wardrobe control
Wearing a tailored oversized camel blazer that looks crisp and expensive, worn open directly against bare skin with no shirt, top, camisole, blouse or garment underneath, no visible fabric inside the lapels, collarbone and neck visible, clean uninterrupted lapel line.

Four synonyms closed off, then the empty space described twice. "No visible fabric inside the lapels" tells the model what those pixels should contain rather than what they should not. The same construction transfers to bare legs, empty backgrounds, hands without jewelry, and any other absence you are trying to specify.

The point Image models are poor at absence and excellent at description. Do not tell it what to leave out. Tell it what is there instead, and name every synonym for the thing you do not want.

How do you turn a look you like into a prompt?

I do not invent a look from nothing. I collect images on Pinterest whose styling I want, and for this model that was a very specific register: oversized taupe and camel tailoring, seamless pale studio backdrops, dewy skin, tousled undone hair, near-shadowless frontal light.

The step that matters is what you do with that image. Use it to produce words, then throw the image away. Upload it once, ask for a description of it, and you get back a paragraph in exactly the vocabulary these models respond to, usually with the palette as hex values. Paste that language into your own prompt.

Do not attach the inspiration image to the generation An attached reference image and a trained identity compete for the same job, because both of them describe a person. Attach a photograph of somebody else and your model's face will drift toward it, which defeats the entire point of training her. Keep the reference field empty and let the trained identity be the only person in the request. The same applies to automatic prompt enhancement: leave it off, or it will rewrite your brief around the reference.

Here is what that looks like in practice. Same trained model, styling language lifted from the reference, nothing attached.

Prompt 3 · a second setup, same face
Editorial beauty portrait of a woman in her late 20s, warm tan skin with natural texture and visible pores, dewy glowing finish, soft warm brown eye makeup with a smudged lash line, thick straight full brows, glossy nude lip, a small beauty mark beside the mouth, bare relaxed expression. Very voluminous dark brunette hair, long and windswept, fine flyaway strands blowing loose across the cheek and mouth. One hand raised to the side of the face, fingers lightly caught in the hair, nails bare. Slim gold hoop earrings. Wearing a plain white ribbed cotton tank, one shoulder visible. Straight-on to camera, chin level, direct unsmiling gaze into the lens, quietly confident. Soft even frontal light, minimal shadow, gentle specular highlight on the skin. Seamless pale seafoam mint backdrop. Shot on 85mm at f2, shallow depth of field, tight head and shoulders crop. Muted warm color grade. Tack sharp on the eyes.
Prompt 4 · full body, seated
Editorial portrait of a woman in her late 20s seated on the floor, one knee drawn up with a forearm resting across it, the other arm lifted overhead with her hand pushing back into the hair at her crown. Warm sunlit tan skin with natural texture, freckles scattered across the nose and cheeks, dewy luminous finish, soft neutral eye makeup, full natural brows, glossy nude lip, lips slightly parted. Long voluminous dark brunette hair sweeping out to one side as if caught in a light breeze, loose waves with flyaway strands. Slim gold hoop earrings. Wearing a plain white ribbed cotton V-neck tank, shoulders and arms bare. Looking up and away off camera past the light, wistful and unposed. Soft directional daylight from the front left, gentle modeling on the shoulders and collarbone. Seamless pale seafoam mint backdrop. Shot on 85mm at f2, shallow depth of field, waist-up seated three-quarter crop. Muted warm color grade. Tack sharp on the eyes.
The AI brand model in a white ribbed tank against a pale seafoam backdrop, hair windswept, hand raised to her hair
Prompt 3Different wardrobe, different world, same woman.
The AI brand model seated on the floor, one arm overhead, looking away from camera
Prompt 4Different pose, different framing, still unmistakably her.

The five-angle set

One image is a picture. A set is an asset. To turn your chosen look into something you can actually use across a site, take the casting prompt and generate five versions of it, and the method is the whole trick: the identity, wardrobe, makeup, hair, lighting and lens language stay identical in all five, and exactly one clause changes.

Swap these five clauses into the opening of the casting prompt. Change nothing else.

VersionThe clause that changesWhat it is for
Straight-on "straight-on and symmetrical to camera with the chin slightly lifted, looking down the nose into the lens, direct unsmiling gaze" The hero. Symmetry reads as authority, and it crops cleanly to square.
Three-quarter "body and head turned in a three-quarter turn to camera right with her eyes coming back to the lens, chin level" The workhorse. Most editorial and most product-adjacent shots live here.
Full profile "in full profile, head turned ninety degrees to camera right, looking straight ahead off frame, clean jawline and nose silhouette against the backdrop" Earrings, necklines, hair.
Wide waist-up "wide waist-up framing with generous negative space above the head and to the sides, full torso and both arms visible ... the boxy oversized cut of the blazer clearly readable" Garment silhouette, and the negative space you need for a headline on a banner.
Extreme close-up "very tight crop running from the collarbone at the bottom edge to the hairline at the top edge, the face filling the frame ... visible pores and fine peach fuzz" Skin, makeup, texture. This is the frame that decides whether a viewer believes she is real.

Add one line to all five that is not in the original: "identical lighting setup." Two words, and they are the reason the set looks like one sitting rather than five sittings.

Variation: the AI brand model straight on to camera with her chin lifted, camel blazer, white backdrop
Version 1 · straight-onChin lifted, looking down the nose into the lens.
Variation: the AI brand model turned three-quarters to camera right with her eyes back to the lens, camel blazer, white backdrop
Version 2 · three-quarterBody turned, eyes returned. Same light, same blazer, same woman.
The prompt is not the skill. Knowing which version to keep is the skill.

Putting her in motion

Once you have a still you are happy with, that still becomes the opening frame of a video. Here is the prompt structure, using the fashion example I ran.

Prompt 5 · the video
A woman crossing the streets of Paris. She's changing her outfits when she crosses the street. Every two seconds, there is an outfit change. Outfit 1 - [describe the first look, or attach the product image] Outfit 2 - [describe the second look, or attach the product image] Outfit 3 - [describe the third look, or attach the product image] Total 8 seconds.

The settings are doing as much work as the words, so set them deliberately.

SettingUse thisWhy
Opening frameYour chosen stillThis is what carries the identity. With no opening frame the model invents a different woman, and your trained face never appears.
Aspect ratio9:16, set in the settingsWriting "TikTok video size" into the prompt text does nothing. Aspect ratio is a setting, not a sentence.
Duration8 secondsEnough for three outfit changes at roughly two seconds each.
Multi-shotOn, automaticThis is what produces the cuts between looks rather than one continuous take.
Prompt enhancementOffSame reason as the stills. Enhancement rewrites your brief.
Do not paste product page links A video model does not browse. A link in a prompt is read as plain text, so what you get back is the model's guess from the words inside the address, not the garment on that page. If the exact piece matters, attach the product images themselves as references, or generate the styled stills first with your model wearing the piece and use one of those as the opening frame. The second route is the one that survives a client review, because you approve the clothes as stills before you spend credits animating them.

What separates this from an amateur AI model?

Not the prompt. Three things, and all three are taste calls rather than tool skills.

  1. Casting for texture instead of beauty. Every line about visible pores, peach fuzz, freckles and no beauty smoothing is doing deliberate work. Smoothing is the default these models reach for and you have to actively suppress it in every prompt. Poreless skin is the fastest way for a viewer to clock an image as generated.
  2. Changing one variable at a time. If you change the hair, the light and the wardrobe in the same round, you cannot tell which change caused what. It looks slower and it is much faster.
  3. Casting for your customer's aspiration, not your own taste. She has to look like the person your customer wants to be, and still belong in the world your product lives in. That is the judgment nobody can hand you in a prompt file, and it is why the same five steps produce a brand asset for one person and a stock photo for another.

The mechanism takes ten minutes. The casting takes as long as it takes. That ratio is the honest summary of this entire discipline.


How do you train an AI model for a brand?

Upload photos of one consistent person, train them into a named Soul Character, and then call that trained identity in every image prompt afterwards. Training runs in the background and takes about ten minutes. The trained identity works with Soul 2.0 and Soul Cinema only, so no other model will accept it.

Why does my AI model's face change between images?

Because something in the request is outranking the trained identity. Attaching a separate inspiration image as a reference, or leaving automatic prompt enhancement switched on, both pull the face toward that reference instead of toward your model. Describe the look in words in your own prompt and leave the reference field empty, and the face holds.

Can an AI video model read a product page link in my prompt?

No. A video model does not open URLs. A link pasted into a prompt is read as text, so the clothing you get back is the model's guess from the words in the address, not the garment on that page. To control the wardrobe, attach the product images themselves as references, or generate the styled stills first and use one as the opening frame.

Can I use a real person's photos to train an AI model?

Only with their explicit permission for this specific use, and not a stranger, a celebrity, or a stock set whose license does not cover AI training, which most do not. Generating the training set instead gives you an identity you own outright with nobody's likeness attached to your brand.

Anne Reyes, brand strategist and founder of Brandworks Studios
About the strategist

Anne Reyes

Brand & AI strategist · Founder, Brandworks Studios

17 years building brands, and I scaled my marketing agency to 7 figures within 2 years, with clients including Amazon, Pier 1, and Dressbarn. I now build AI marketing systems for e-commerce brands through BAIBE, the Brandworks AI Brand Engine. AI does the production. Taste directs it.

Keep going

Want the next build?

The Chic Strategist breaks down a new AI marketing system every week. Or bring your brand to a strategy call and we'll build it together.

← All the free guides