# Photo Edit Brief https://photo-edit-brief.skillsafe.ai/ Tells you whether a requested photo edit can survive a re-render, translates every "keep this unchanged" into words a renderer can act on, writes the labelled edit brief from your photograph, renders it, and compares the result against your original. ## The problem it exists for There is no image-in / image-out model on this platform. An image-generation run refuses file attachments outright. So the model that paints the picture never sees the photograph — it reads words and paints something new. That matters because the source skill's entire Edit half is written as invariants: "preserve face, body shape, pose", "keep the background unchanged", "preserve label text exactly". An invariant is a promise about pixels that already exist. A renderer with no pixels cannot keep one by negation. It can only keep it by description — repainting the thing closely enough that it reads as unchanged. That converts cleanly for some edits and not at all for others, and which is which is not obvious. Working it out before you spend anything is what this app is for. ## The survival matrix Eight edit types, each rated on what description can actually carry. The rating comes from the invariant clause the source skill uses to define that edit. - **lighting-weather** — survives. The whole edit is atmosphere, and atmosphere is what words carry best. Holds for places and objects; not for a specific person. - **sketch-to-render** — survives. Layout, proportion and perspective are geometric facts a sentence states exactly, and a drawing has no photographic identity to lose. - **style-transfer** — survives. Never really an edit: the reference is a mood board, and palette, texture and brushwork describe well. - **text-localization** — degrades. Layout and replacement copy carry; logos and photographic content inside the frame are re-invented. - **precise-object-edit** — degrades. Camera angle and lighting carry; "surrounding objects unchanged" does not, because they get repainted from the description. - **compositing** — degrades. The relationship between the elements carries; the specific subject and specific background do not. - **background-extraction** — impossible. A cutout is defined by keeping your exact pixels. Use a client-side segmentation tool instead; this app names two. - **identity-preserve** — impossible. Every word you could write about a face fits thousands of faces. No path exists from a photograph to the renderer, so this app cannot face-swap. Two conditions move a rating: a specific real person in the frame drags several edits to impossible; ordinary rather than particular surroundings lifts the degrading ones. ## The three lanes - `task: "brief"` — a text run on `gpt-terra` (`gpt-5.6-terra`) WITH the photograph attached. The writer reads the picture and returns one JSON object: `title`, `seen` (what it actually saw, shown to you so you can catch a misreading before spending), `spec` (thirteen labelled lines), `translated` (each preserve-clause and the description it became), `unsupported` (what this brief cannot deliver), `why`. Or `{"refused": true, "reason": "..."}`. - `$model: "gpt-image"` — an image run on `gpt-image`. Exactly two keys, `instruction` and `$model`; attachments are refused here by the platform. Output is `job.output.images[0].b64`. - `task: "check"` — a text run on `gpt-terra` with BOTH pictures attached. Returns `verdict` (close / partial / off), `checks` (per item: held / drifted / lost), `summary`, and `next_change` — exactly one targeted change, because iterating with one change is what converges. ## The prompt schema The source skill's labelled spec, with one deliberate omission. The source carries an `Input images:` line ("Image 1: edit target"). That line cannot exist here — it names an attachment the renderer is structurally unable to open, so it is a pointer by definition. Dropping it leaves thirteen lines: Use case, Asset type, Primary request, Scene/backdrop, Subject, Style/medium, Composition/framing, Lighting/mood, Color palette, Materials/textures, Text (verbatim), Constraints, Avoid. Seven are required: use case, primary request, subject, style, composition, lighting, constraints. ## Languages Eleven: English, Simplified Chinese, Japanese, Korean, Spanish, Brazilian Portuguese, French, German, Russian, Indonesian, Vietnamese. **There is one language control, and it decides both things.** The picker in the header sets the interface language AND the language the brief is written in. There is no separate "output language" setting and, deliberately, no detection of the language you typed your request in — a heuristic over your prose is a second control that silently disagrees with the first. The model is told the language by NAME ("Japanese"), not by code, and the writing guide has an explicit section acting on that field. Four things never change with it: the reply's JSON keys, the `use_case` slug, the quoted half of each invariant translation, and any words you asked to be lettered into the picture — those are painted as shapes rather than read as language, so they are reproduced exactly as you wrote them. Two honest limits. The request classifier's evidence terms are English, so in another language it simply does not guess and the picker is used directly — the verdict comes from your pick either way. And the preserve-clause detector carries a small, high-precision marker set per language; it is tuned to miss rather than to fire wrongly, because in a finished brief a hit blocks the render. ## What is free The survival verdict, the request classifier, the invariant translator, the thirteen-line linter, the worked examples, and every export. All of it runs in your browser. Only the three model runs cost credits. The linter blocks a render — rather than warning — on four things, because each one wastes a paid run that a free rewrite would have fixed: an unfilled bracketed placeholder, a phrase addressed to a chat assistant, a phrase pointing at the attachment, and any surviving preserve-clause. ## Privacy Photographs are resized and re-encoded through a canvas in your browser before upload, which writes a fresh file from pixels and so drops every EXIF block including location. The brief lane and the check lane upload the re-encoded copy because the writer has to see it. The render lane uploads nothing. ## Sources The use-case taxonomy (sixteen exact slugs), the labelled prompt schema and the per-slug prompting rules come from the `imagegen` skill by @openai, https://github.com/openai/skills. This app is a derived work: the survival matrix, the invariant translator, the linter, the check lane and the interface are original. ## Pages - / — the app - /api.html — driving the app programmatically, with samples in eight programming languages - /tokens.html — token management (noindex) - /guide.js — the exact writing guide this app sends the model - /slugs.js — the taxonomy and the survival matrix as data - /lang/en.js — the UI strings; one file per locale beside it