Set up AI descriptions
This page is for administrators. It covers turning on AI descriptions, choosing a model, tuning the prompt and backfilling an existing library. The settings live in Administration, then Settings, then Machine Learning Settings, then Image descriptions and tags.
Before you start
Section titled “Before you start”- Your machine learning container is working, ideally with hardware acceleration. See Hardware acceleration.
- The model you pick is a Qwen2.5-VL or Phi-3.5-vision model, not Florence-2.
- For names in descriptions, facial recognition is on and you’ve named some people. See People.
- Videos need nothing extra. Each video’s moment frames are cut the first time it’s described.
Setup order
Section titled “Setup order”| Step | Do | Why now |
|---|---|---|
| 1 | Pick the right model | Florence-2 ignores every other setting on this page. Get the model right before you tune anything. |
| 2 | Run facial recognition and name your most-photographed people | Descriptions only use names you’ve given. Doing this first means the first run already says “Kelly” instead of “Someone”. |
| 3 | Preview a description on a few photos and videos | Nothing is written, so you see the result before paying for a library-wide run. |
| 4 | Tune the prompt | Tuning before the big run means you don’t describe everything twice. |
| 5 | Turn on identity injection | Cheap to change; set it together with the prompt. |
| 6 | Re-queue all descriptions | One library-wide pass with everything set the way you want. |
| 7 | Turn on smart albums | Smart albums fill from description tags, so they need descriptions first. |
| 8 | Re-evaluate smart albums once descriptions are done | A one-time backfill for older photos. |
If you’ve already started without following this order, that’s fine: you can re-queue at any time, and smart albums can be re-evaluated one album at a time.
Quick setup
Section titled “Quick setup”- Go to Administration, then Settings, then Machine Learning Settings.
- Under Image enrichment hardware, choose Auto-detect, Intel iGPU (OpenVINO) or NVIDIA GPU (CUDA).
- In Image descriptions and tags, make sure Generate image descriptions and tags is on.
- Choose a Description model. The default,
Qwen/Qwen2.5-VL-3B-Instruct, suits most servers. - Open Prompt & Vocabulary, then Identity injection, and make sure Enable identity injection is on.
- Save.
- Use Preview a description to check the result on a few of your photos and videos.
- Under Status & Re-generation, choose Re-queue all image descriptions, check the estimate, and confirm.
- Turn on smart albums if you want them.
New uploads are described automatically after their thumbnails are made, using your current settings.
Choose a model
Section titled “Choose a model”| Hardware | Model | Notes |
|---|---|---|
| NVIDIA with 16 GB or more video memory | Qwen/Qwen2.5-VL-7B-Instruct |
Best quality of the common choices |
| NVIDIA with 6 to 12 GB | Qwen/Qwen2.5-VL-3B-Instruct |
The default; a good balance |
| Intel integrated graphics | Qwen/Qwen2.5-VL-3B-Instruct |
Runs as llmware/qwen2.5-vl-3b-ov through OpenVINO |
| CPU or small integrated GPU | microsoft/Phi-3.5-vision-instruct |
Lighter, still follows the prompt |
| Older Phi build | microsoft/Phi-3-vision-128k-instruct |
Smaller, slightly lower quality |
| Last resort | microsoft/Florence-2-base-ft |
Captions only |
The Description model list shows an estimate of the video memory each model needs:
| Model | Video memory | Notes |
|---|---|---|
| Qwen2.5-VL 3B (default) | About 6 GB | Good at objects and scenes. On OpenVINO it runs as llmware/qwen2.5-vl-3b-ov. |
| Qwen2.5-VL 7B | About 16 GB | Better at composition, text in images and subtle scenes. On OpenVINO it runs as llmware/qwen2.5-vl-7b-ov. |
| Qwen2.5-VL 32B | About 64 GB | Complex scenes and fine detail. NVIDIA only. |
| Qwen2.5-VL 72B | About 144 GB | Best quality in the family; needs several GPUs. NVIDIA only. |
| Qwen3-VL 30B-A3B | About 60 GB | Only about 3B parameters active at a time, so it runs near 7B speed. NVIDIA only. |
| Phi-3.5-vision-instruct | About 5 GB | Smaller alternative on OpenVINO |
| Phi-3-vision-128k | About 5 GB | Older and smaller still |
| Florence-2-base-ft | About 1 GB | Captions only; local fallback |
| Florence-2-large-ft | About 3 GB | Captions only; local fallback |
Custom… lets you type any Hugging Face model ID, but only the Qwen2.5-VL, Qwen3-VL, Phi-3 and Phi-3.5-vision, and Florence-2 families load. Anything else fails with Failed to load model in the worker logs.
A model downloads the first time a job uses it, which can take several minutes.
Fallback model
Section titled “Fallback model”Fallback model is used on this server when the main model can’t run. Florence-2 is the usual choice because it’s small enough to share a GPU with other models. If the main model fails on a local machine learning server, Frameleaf retries the same request with the fallback. Frameleaf Cloud always uses its chosen model, with no fallback.
Hardware
Section titled “Hardware”On Intel, use the OpenVINO machine learning image and leave the description device on AUTO, which picks the best device and falls back when the GPU isn’t available. On NVIDIA, use the CUDA image and choose NVIDIA GPU (CUDA).
Description models download from huggingface.co unless you set HF_ENDPOINT to your own mirror. See Where models come from. Your photos are processed by your own machine learning container, and Frameleaf sends no telemetry.
Preview before you commit
Section titled “Preview before you commit”Preview a description opens Try an enrichment change:
-
Choose sample. Pick up to six of your own photos or videos, and where to run them. The destination is always named; nothing falls back to another one.
-
Compare. The model and prompt you’re editing, saved or not, run on each sample. The current description and the new one appear side by side, with the tags, any names the check removed and, for a video, how many frames it saw. Nothing is written: descriptions, tags, the Locked state, search data and frames stay as they are.
-
Scope. Choose the stages a background plan should run on the samples:
- Descriptions and tags, and the Locked-content check, for photos
- Reusable video frames, the moment search index and moment captions, for videos
The moment stages need the frames, so ticking one ticks the frames too. Moment captions are never ticked for you: they add one model request per frame, and the dialog says how many.
A plan uses the saved model and prompt, and records them on every result it writes. It runs in the background one item at a time, shows in Activity, and survives closing the browser or restarting the server. You can Pause or Cancel it between items. Each failure is retried once automatically, and Retry starts a new plan over only the items and stages that didn’t finish.
Prompt & Vocabulary
Section titled “Prompt & Vocabulary”| Setting | What it does | Default |
|---|---|---|
| Description style | Terse (one or two sentences), Balanced (a short paragraph) or Rich (a detailed description) | Balanced |
| Sentence count target | Target number of sentences, from 1 to 6. The model sometimes goes one over. | 3 |
| Look for | Categories the model should point out when they’re visible | brands, signage, screens, documents, uniforms, tools, vehicles, animals, food, landmarks |
| Custom vocabulary | Tag values the model should reuse, spelled exactly as you write them | Empty |
| Custom instructions | Your own guidance in full sentences, up to 2,000 characters | Empty |
| Forbidden inferences | Things the model must not infer, even when the photo suggests them | diagnoses, medication names, procedures, pregnancy, disability |
| NSFW indicators | Explicit terms allowed in descriptions of sensitive photos. Clear the box and save to restore the defaults. | A built-in list |
| Medical indicators | Medical terms allowed in descriptions. Clear the box and save to restore the defaults. | A built-in list |
| Identity injection | Names recognised people in descriptions | On, 5 names, 0.7 |
| Advanced (raw prompt editor) | Replaces the whole prompt with your own template | Off |
List settings take one entry per line.
Settings by library type
Section titled “Settings by library type”| Library | Style | Identity injection | Add to Look for | Add to Custom vocabulary |
|---|---|---|---|---|
| Family | Balanced or Rich | On | birthday cake, sparklers, candles, prom, graduation, recital, soccer ball, baseball glove | golden hour, overcast, dappled light, backyard, park, beach |
| Documents and receipts | Terse | Off | total amount, store name, line items, payment method, invoice number, due date, expiry date, signature line | Keep it short |
| Travel | Balanced | On | mountain range, beach, hotel lobby, train station, food market, street scene, landmark | The landmarks and cities you visit often |
| Pets | Terse or Balanced | Either | breed, leash, food bowl, harness, fetch, asleep, swimming, kennel | Your pets’ breeds, such as golden retriever or tabby |
For a documents library, also add account numbers verbatim, full credit card numbers, social security numbers to Forbidden inferences.
For a hobby, teach the model its words through Custom vocabulary: long exposure, leading lines, bokeh for photography, latte art, sourdough, plating for cooking, or gravel, derailleur, peloton for cycling.
Custom instructions
Section titled “Custom instructions”Custom instructions is the easiest way to change how the model writes without touching the whole prompt. Treat it like a short brief to a careful assistant: write full sentences, say what to do and when not to, and keep it to a couple of paragraphs. The text is added to the prompt before the output rules, so the model reads it first.
Some examples that work well:
Vehicles
If you can clearly see a car, truck, or motorcycle, identify the make and model inthe description (e.g. "a red Tesla Model 3", not "a red car"). If you are uncertainabout either the make or the model, just say "a red car" rather than guessing.Sports
When people are clearly playing a sport, name the sport in the description (soccer,baseball, basketball, tennis, etc.) and mention the visible equipment they're using.Do not guess the sport from clothing alone, only from visible play, equipment, orfield markings.Travel
When the photo is clearly outdoors and looks like travel, identify the location type:beach, mountain, forest, city street, market, train station, airport, museum, cathedral,temple, ruins, hotel lobby. If a recognizable landmark is visible, name it. Do notinvent landmarks if you are not sure.Documents and receipts
When the photo is a document, receipt, or screenshot of text, transcribe the mostimportant visible identifiers: store name, total amount, date, invoice number, anddocument type. Do not transcribe full account numbers, full credit card numbers,social security numbers, or any other sensitive personal data even if visible; referto them generically (e.g. "account number redacted").Food
For food photos: identify the dish by name if recognizable, name the cuisine if youcan tell, and note the plating style (rustic, fine-dining, casual, street-food). Donot name a dish you are not confident about; say "pasta dish" rather than guessing.Pets
For pets: identify the breed when you can recognize it. For dogs in particular, notethe activity (sleeping, playing fetch, swimming, on a leash, in a car). Do not guessa breed from coat color alone.Tone
Keep descriptions plain and factual. Do not use poetic language ("a tapestry ofcolors", "bathed in golden light"). Do not editorialize about the people or events("a joyful family", "an intimate moment"). Stick to what is visible.A family library, all in one
Always name every recognized person and avoid generic group terms like "the family"or "everyone".
If you see a car, name the make and model when recognizable.If people are playing a sport, name the sport and any visible equipment.For pets, identify the breed when recognizable.For travel scenes, name the landmark if you are certain.
Keep descriptions factual and short. Do not use poetic language.What custom instructions can’t do
Section titled “What custom instructions can’t do”- They can’t override the safety rules. The sensitive-content and medical handling, and the check that removes invented names, still apply.
- They can’t add names. A name that isn’t recognised on the photo is still replaced with “Someone”.
- They can’t change the output format. Use the raw prompt editor for that.
- They can’t make the model remember anything between photos. Each one is described on its own.
- They’re ignored while the raw prompt editor is on.
Identity injection
Section titled “Identity injection”| Setting | What it does | Default |
|---|---|---|
| Enable identity injection | Passes the named people in each photo to the model | On |
| Max names | How many people to name in one photo, from 1 to 20 | 5 |
| Min face confidence | The confidence a face match needs before its name is used, from 0 to 1 | 0.7 |
Lower Max names to 1 or 2 for crowd photos where only the main people matter. Raise it to 10 to 15 for teams, weddings and other big groups. If a photo has more named people than Max names, some are left out.
Raw prompt editor
Section titled “Raw prompt editor”Most people should leave the Advanced raw prompt editor alone. The other settings build the same prompt with safer defaults. Use it only when you need something they can’t express, such as a different output format or sections in another order.
- When you turn on Enable raw prompt override with an empty box, Raw prompt template fills with the current default template, so you start from something that works. It never overwrites text you’ve already written.
- Reset to default replaces your template with the default, discarding your edits.
- Turning the override off goes back to the structured settings. Your template is kept for next time.
You can use these placeholders:
| Placeholder | Becomes |
|---|---|
{names} |
The recognised people, with the wording that requires the model to use them. Empty when identity injection is off or nobody is recognised. |
{schema} |
The output format the model must return. Required in strict mode. |
{vocabulary} |
Your custom vocabulary |
{style_hint} |
The length and tone cue for the chosen style |
Placeholder validation is either Strict, where saving fails without {schema}, or Warn, where it saves and shows a warning. Strict is the default.
Custom instructions aren’t added to a raw template; paste the text where you want it. Frameleaf still adds the video timing note for videos, and extra care for photos flagged as sensitive, around your template.
If the template is wrong, descriptions come out empty or broken. Turn the override off to recover.
Describe your library
Section titled “Describe your library”Status & Re-generation shows:
| Field | Shows |
|---|---|
| Last config change | When the description settings last really changed |
| Pending re-queue scheduled | Set when you chose Re-queue later |
| Total eligible image assets | Everything descriptions can process |
| Already described | Items with a description |
| Pending re-description | Items without one |
| Estimated re-queue time | Based on the last 100 finished description jobs |
Re-queue all image descriptions opens a window with the counts, the model, the seconds per item and the estimated total time:
- Cancel closes it.
- Re-queue later leaves a reminder banner at the top of the settings, so you can make several changes and run once at the end. Re-queue now on the banner starts it.
- Re-queue now starts straight away. Clicking twice doesn’t queue the work twice.
When descriptions go to Frameleaf Cloud, nothing is queued here; Frameleaf Cloud describes photos in batches instead.
After a restart, the estimate uses 1.5 seconds per item until 100 jobs have finished. That’s about right for Qwen2.5-VL 3B on a typical Intel integrated GPU. On an NVIDIA GPU, expect around 0.5 to 1 second per item with the 3B model.
You can also run Image descriptions and tags, then All, from Administration, then Jobs. Jobs skip items that already have a description unless you force them, and a forced run never adds the AI description: block twice.
Describe video moments
Section titled “Describe video moments”Describe video moments, under the Descriptions & tags routing settings, is off by default. When it’s on, every video described automatically also gets its frames captioned straight afterwards, one extra model request per frame. It doesn’t change any description, so turning it on or off never asks for descriptions to be redone. Frames looked at per video beside it shows the sampling policy, six frames, and can’t be changed.
Sensitive-content detection
Section titled “Sensitive-content detection”If Detect NSFW images is also on, Frameleaf checks each photo for sensitive content first and passes the result to the description model, so the description stays factual. See Private content.
A worked example
Section titled “A worked example”Say you have about 80,000 photos and 3,000 videos, and one NVIDIA GPU with 8 GB of memory.
- Day 1. Turn on facial recognition, run face detection and wait for it to finish. Name your 20 or so most-photographed people; the People page shows them first.
- Day 2. Choose
Qwen/Qwen2.5-VL-3B-Instruct, which fits comfortably in 8 GB. Set Balanced and 3 sentences, add a few things to Look for and Custom vocabulary, paste the family example into Custom instructions, and set Max names to 8. Save, then preview on a handful of photos and videos. - Day 3. Re-queue all descriptions. At 0.5 to 1 second each, expect 12 to 24 hours. Turn on smart albums, add any extra triggers, and click Re-evaluate all assets when descriptions finish; that takes minutes.
- Day 4. Open ten random photos and videos. Are names used? Are your instructions followed? Do smart albums make sense? Do video descriptions use words such as “begins” and “then”? Adjust, rerun a few items from their info panels, and repeat until you’re happy.
- After that. New uploads are described automatically. New names appear in new descriptions straight away; rerun older ones if you want them updated.
Troubleshooting
Section titled “Troubleshooting”Descriptions are still generic with identity injection on
Section titled “Descriptions are still generic with identity injection on”- Check the photo has at least one named face in its People list.
- Check the model is Qwen or Phi, not Florence.
- Check Max names is at least the number of named people in the photo.
- Rerun descriptions and tags for the photo from the Image enrichment section of its info panel. Older descriptions keep the wording they were written with.
Custom instructions aren’t working
Section titled “Custom instructions aren’t working”- Check the raw prompt editor is off.
- Check you saved.
- Rerun some photos. Existing descriptions don’t change by themselves.
- Try plainer, more direct wording: “If you see a car, identify the make and model” works better than “try to be more specific about vehicles”.
- Check you’re under 2,000 characters; longer text can’t be saved.
A video shows “video-frames-unavailable”
Section titled “A video shows “video-frames-unavailable””No frames could be cut. The video is very short, longer than three hours, or unreadable. Choose Find moments in its Moments section to try again. If that fails, check that video processing works for other files and that the file isn’t damaged.
The raw prompt box is empty
Section titled “The raw prompt box is empty”Turn the override off and on again; it fills itself only when the box is empty. Or click Reset to default.
The reminder banner won’t go away
Section titled “The reminder banner won’t go away”It clears when a re-queue starts. If the job failed, check Administration, then Jobs, then click Re-queue now on the banner.
The estimate looks wrong
Section titled “The estimate looks wrong”After a restart the estimate uses a default until 100 jobs have finished, then settles over the next 100 or so.
A real name was replaced with “Someone”
Section titled “A real name was replaced with “Someone””The person isn’t named on that photo. Name them in its People list and rerun it, or edit the description by hand. There’s no separate list of allowed names.