Skip to content

AI descriptions

Frameleaf can describe your photos and videos in plain sentences and tag what’s in them, all on your own server. The descriptions and tags are searchable, can name the people you’ve recognised, and feed smart albums that gather your travel, food and pet photos automatically.

Your administrator controls whether descriptions run and how they’re written. The setup is covered in Set up AI descriptions.

  • An AI description added to each photo’s description field, in an AI description: block. Anything you wrote yourself is kept: the generated block is added once, after your text, and never replaces it.
  • Searchable tags for visible objects, scenes, text and context, all lowercase and without duplicates. They’re ordinary tags, so you can browse and filter by them.
  • Names, not “a family”. With identity injection on, a description says “Kelly and Connor playing baseball” instead of “a family playing baseball”.
  • Video descriptions that cover the whole clip, not just its first frame.

For example, a photo of a car might be described as “A red Tesla Model 3 parked in a driveway at dusk”, with tags such as car, driveway and dusk.

Descriptions make your library easier to search in three ways:

  • Smart search gives a boost to photos whose description matches your words.
  • The Description search mode finds words in descriptions, yours and generated ones.
  • Tags from descriptions work with the tag: filter and the Tags filter.

See Search and Search syntax.

When identity injection is on, Frameleaf tells the model who it recognised in the photo, using the names you’ve given in People, and where they are in the frame. The model has to use each name at least once and mustn’t fall back on words such as “a family”, “a group” or “everyone”.

Frameleaf then checks the result:

  • Only your names are used. Any other name the model writes is replaced with “Someone”. Days, months, holidays and well-known places, such as “Monday”, “December”, “Christmas” or “Paris”, are allowed.
  • One person, named. If one known person is in the photo and the model writes “the woman” or “the boy”, Frameleaf puts the name in.
  • Several people. If several known people are in the photo and the model uses a general word, Frameleaf can’t know who was meant, so it leaves the sentence alone.

Only people with a name, who aren’t hidden, are passed to the model. Faces that were detected but not named are skipped.

Fixing a face or a name in People marks the descriptions that used the old name as out of date, and they’re redone the next time descriptions are processed. Naming someone new only affects descriptions written after that; your administrator can redo older ones.

If a description replaced a real name with “Someone”, the person probably isn’t named on that photo. Add them in the photo’s People list and ask your administrator to rerun the description, or edit the description yourself.

Frameleaf cuts six evenly spaced frames from each video, ranks them, and keeps them as the video’s moment frames. To describe the video, it lays the frames out as one grid, in time order, and tells the model how long the video is and when each frame was taken. The model describes the video as a whole, including what changes from frame to frame.

A video description might read: “A short clip of Connor swinging a bat in a backyard; he begins facing the camera, swings the bat to his right, and walks off-frame at the end.”

The faces Frameleaf found in the video’s thumbnail are used to name the people in it.

A video that’s shorter than a fraction of a second, longer than three hours, or can’t be read gets no description. Its status shows skipped with the reason video-frames-unavailable. Open its Moments section and choose Find moments to try again.

If the part of a video you care about falls between the six frames, add your own moment at that time. Your moments are searchable and are never replaced.

A video’s Moments section in the info panel shows its frames ranked best first, the cover frame and its timestamped moments. Choose a frame or moment to play from there.

  • Cover: pick any frame as the cover. It’s kept as a time in the video, so it survives the frames being cut again. Use the best frame goes back to the best-ranked one.
  • Your moments: add a moment at any time with a title and an optional typed transcript. There’s no automatic speech recognition. Your moments are never removed by a refresh, a replaced original or a face correction.
  • Find moments or Refresh moments cuts the frames and builds the moment search index. Add captions describes each frame on its own, which takes one extra model request per frame.
  • Find similar looks for moments like the one you’re watching in your other videos.
  • Search: searches show Moments in your videos above the results, matching frames by meaning and your moments by their words. Only your own videos are searched.

Each generated result remembers what it was made from. When one of these changes, only the results that depend on it are affected:

Change What happens
The original file is replaced Frames, their search data and generated moments are removed. The description is marked out of date and is never shown against the new file.
A face correction changes the names in a photo Generated descriptions and captions that used the old names are marked out of date and redone next time.
Your administrator changes the prompt Generated descriptions are marked out of date.
Your administrator changes the smart search model Only the frames’ search data is cleared. Frames and moments stay.

Descriptions you wrote, your own moments, typed transcripts and your cover choice are never touched.

The description is an ordinary field: open a photo’s info panel and edit it as you like.

Administrators also see an Image enrichment section in the info panel, with the model, status and any error. From there they can:

  • rerun descriptions and tags for that photo
  • clear the generated AI description: block without touching your own text
  • clear the generated tags from the photo, without deleting the tags themselves

Tags from descriptions feed six ready-made albums: Travel, Documents & Receipts, Screenshots, Food, Pets and Nature. They fill themselves as descriptions finish, and anything you take out stays out. See Built-in smart albums.

  • Descriptions are made by your own server’s machine learning container, or by a computer on your home network that your administrator has added. Your photos only go to Frameleaf Cloud if your administrator routes descriptions there and records consent.
  • Locked photos are described like any other, but their descriptions and tags only appear in your unlocked session.
  • Tags are visible to anyone you share a photo with. Don’t rely on a tag or description to keep something private.

Descriptions are set up in Administration, then Settings, then Machine Learning Settings, then Image descriptions and tags. The full guide is Set up AI descriptions.

  1. Pick a Qwen or Phi model. Florence-2 ignores the prompt settings.
  2. Name your most-photographed people in People.
  3. Preview a description on a few photos and videos. Nothing is saved.
  4. Tune the prompt and add custom instructions.
  5. Turn on identity injection.
  6. Re-queue all descriptions.
  7. Turn on smart albums, then re-evaluate them once descriptions finish.

Doing it in this order means you only describe your library once. Each step is explained in Set up AI descriptions.