← Photo Edit Brief API
Get a token

Driving Photo Edit Brief over HTTP

Everything the page does, you can do directly. Base URL: https://api.skillsafe.ai/v1/app-api. Every response is {"ok":true,"data":{…}} or {"ok":false,"error":{…}}.

The route shape is not what it looks like. There is no /apps/{slug}/ segment anywhere. The endpoints are /guest, /me, /files, /estimate, /run, /run-stream and /jobs/{id}, and the slug is bound to your token when you mint it. Getting this wrong returns 404 not_found.
An image run refuses attachments. POST /run with $model: "gpt-image" plus $files returns 400 validation_error. There is no image-in / image-out call on this platform, which is the whole reason this app is shaped the way it is: a text model reads your photograph and writes a description, and the image model paints from that description alone. Note also that /estimate happily approves the body that /run rejects — a clean estimate is not evidence a run works.

The three lanes

Two text lanes routed on a task field, plus one image lane selected by a $model override. Every lane also takes guide, whose value is the full text of /guide.js — the writing instructions live in a served asset rather than the system prompt, because on an image run the system prompt is joined to the input and anything instruction-shaped in it gets painted into the picture as words.

LaneSelected byModelAttachmentsReturns
brief"task": "brief"gpt-terra → gpt-5.6-terrathe photograph, via $filesone JSON object of spec lines
render"$model": "gpt-image"gpt-imagenone — refusedjob.output.images[0].b64
check"task": "check"gpt-terra → gpt-5.6-terraboth pictures, original firstone JSON report object

The language field

Both text lanes take language — the language the model must WRITE in, spelled out as a name rather than a code: "Japanese", "Simplified Chinese", "Spanish". A companion language_code carries the short form (ja, zh) for your own bookkeeping; the model acts on the name.

The guide reads this field. That is worth stating because the failure it avoids is silent: a language field that the prompt never mentions is tokens you pay for and the model ignores, and the output comes back in whatever language it felt like. The guide at /guide.js has an explicit section acting on it, and names the four things that do NOT change with it — the JSON keys, the use_case slug, the quoted half of each invariant translation, and any words destined to be lettered into the picture.

Accepted names, matching the app's own picker: English, Simplified Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Russian, Indonesian, Vietnamese. Omit the field and you get English.

1. A tiny client

Error handling on the envelope, once, so the rest of the steps stay readable.

2. Get a token

A guest token reads the app and prices a run. The model runs need a signed-in account, because each one spends credits — /tokens.html hands you a personal one without opening DevTools.

3. Who am I, and can I afford it

GET /me returns exactly three fields: subject_type ("user" or "guest"), subject_id and credits. Signed-in is subject_type === "user" — there is no email and no name to test.

4. Upload the photograph

POST /files is multipart and returns a file_id. Two things worth knowing before you build against it: file ids are scoped to the subject that minted them, so an id created as a guest returns 404 the moment you sign in; and /estimate both prices attachments (about +35 credits each) and validates that they exist, so a body that omits them under-quotes the run and one carrying a stale id fails to price at all.

5. Price it — free, and it creates no job

POST /estimate takes the same body as /run and charges nothing. Assert three things on the response: model is gpt-5.6-terra, model_alias is gpt-terra, and markup_bps is 1000. Present hold_credits as reserved, never as the price — the hold prices the full output cap and the actual charged_credits is usually far lower.

6. Lane one — write the brief

A text run with the photograph attached. Pass an Idempotency-Key header on every run, derived from a hash of the body plus an attempt counter: without one, a network retry can bill twice.

What comes back

One JSON object on job.output.output, as a string you parse yourself:

{
  "title":  "three to six words",
  "seen":   "what the writer actually saw in your photograph",
  "spec":   { …the thirteen labelled lines… },
  "translated":  [ { "kept": "keep the background", "as": "the description it became" } ],
  "unsupported": [ "what this brief cannot deliver" ],
  "why":    "one sentence"
}

Or, if the request is refused: {"refused": true, "reason": "…"}.

7. Lane two — render it

Exactly two keys. Every additional key is concatenated into the text the image model sees and painted as literal words, so this body must carry nothing but the prompt and the model override. Output arrives at job.output.images[0].b64; job.output.output is the empty string.

8. Lane three — check what survived

A text run with both pictures attached, original first. Returns verdict (close / partial / off), checks (each held / drifted / lost), summary, and next_change — one targeted change, because iterating with a single change is what converges.

9. Streaming

POST /run-stream returns Server-Sent Events for the two text lanes. Read the SSE frames yourself as below; note that in a browser the SDK's onDelta callback receives keep-alive ticks rather than text deltas, so a progress bar built on it will never move.

Error codes

CodeHTTPWhat it means
not_found404Wrong route. The app-API endpoints have no /apps/{slug}/ segment — the slug is bound to the token at /guest.
validation_error400The body was rejected. On an image run this is usually $files: attachments are not supported there at all.
unauthorized401No token, or a stale one. Mint a fresh one; guest tokens expire.
insufficient_credits402The balance is below min_credits. Call /estimate first and check against /me.
rate_limited429Back off and retry. Do not tight-loop a poll.
internal500A platform-side failure. Retry once with the SAME idempotency key so a completed job is not re-billed.

The use-case taxonomy

Sixteen exact slugs from the imagegen skill by @openai. The eight Edit slugs carry a survival verdict; the invariant column is the clause the source skill defines that edit by, and is exactly what a renderer with no pixels cannot keep by negation.

SlugEditSurvives a re-render?The invariant it is defined by
lighting-weatherLighting & weathersurvivespreserve subject identity, geometry, camera angle, and composition; change only lighting, atmosphere, and weather
sketch-to-renderSketch to rendersurvivespreserve layout, proportions, and perspective; choose realistic materials and lighting; do not add new elements or text
style-transferStyle transfersurvivespreserve palette, texture, and brushwork; no extra elements
text-localizationText & localizationdegradeschange only the text; preserve layout, typography, spacing, and hierarchy; no extra words; do not alter logos or imagery
precise-object-editObject add / remove / replacedegradespreserve camera angle, room lighting, floor shadows, and surrounding objects; keep all other aspects unchanged
compositingCompositingdegradesmatch lighting, perspective, and scale; keep the base framing unchanged
background-extractionCutout / transparent backgroundimpossiblecrisp silhouette; no halos or fringing; preserve label text exactly; no restyling
identity-preserveIdentity-preserving editimpossiblepreserve face, body shape, pose, hair, expression, and identity; match lighting and shadows

The eight Generate slugs have no original to preserve, so the re-render gap does not apply.

SlugKind
photorealistic-naturalPhotorealistic / natural
product-mockupProduct mockup
ui-mockupUI mockup
infographic-diagramInfographic / diagram
logo-brandLogo / brand mark
illustration-storyIllustration / story
stylized-conceptStylized concept
historical-sceneHistorical scene

The prompt schema

Thirteen labelled lines, joined as Label: value in this order and sent to the image model. The source skill's schema has a fourteenth, Input images: — it is omitted here because it names an attachment the renderer cannot open, which makes it a pointer by definition.

KeyLabelWhat goes in it
use_caseUse caserequiredOne of the sixteen taxonomy slugs.
asset_typeAsset typeoptionalWhere the picture will be used.
primary_requestPrimary requestrequiredThe change, in one sentence.
sceneScene/backdropoptionalThe environment, described as if new.
subjectSubjectrequiredThe main subject, concretely.
styleStyle/mediumrequiredPhoto, illustration, 3D - and the register.
compositionComposition/framingrequiredCamera height, distance, what sits where.
lightingLighting/moodrequiredDirection, hardness, colour of the light.
paletteColor paletteoptionalThe actual colours and their relationship.
materialsMaterials/texturesoptionalReal surfaces and how they catch light.
textText (verbatim)optionalEvery string quoted letter-for-letter.
constraintsConstraintsrequiredWhat the picture must do. Positive, not preserve-clauses.
avoidAvoidoptionalNegative constraints - things that must not appear.

The app's own linter blocks a render on four things, each of which would otherwise waste a paid run: an unfilled bracketed placeholder, a phrase addressed to a chat assistant, a phrase pointing at the attachment, and any surviving preserve-clause. It warns above about 1800 tokens.