Image to JSON Prompt: A Reference Workflow You Can Test
A model-agnostic workflow for extracting visible evidence from a reference image, editing it field by field, generating variants, and checking results without promising an exact recreation.
Contents

An image-to-JSON prompt is most useful when you treat it as an editable visual specification—not as a way to recover the secret prompt that created an image. The practical workflow is: describe only what is visible, separate fixed details from changes, generate several candidates, and compare each candidate with the reference using the same checklist.
That distinction matters. A JSON object can make subject, composition, lighting, materials, and text easier to edit, but it cannot reveal an unseen camera model, a hidden light source, an exact font name, or the original creator’s wording. It also cannot guarantee a pixel-for-pixel copy.
This guide gives you a copy-ready analysis prompt, a clean JSON schema, an editing method, and a pass/fail review process.
What this workflow is—and is not
Use this workflow when you want to:
- preserve the visual logic of a reference while changing selected details;
- compare several image models without rewriting a long prompt each time;
- diagnose why a generated result drifted;
- build a repeatable review process for photos, ads, posters, interfaces, and illustrations.
Do not use it to claim that you have recovered the original prompt. A reference image is an output, not a complete record of how it was made.
A useful JSON prompt has three properties:
- Observable: every factual statement can be supported by the pixels.
- Editable: important visual dimensions have separate fields.
- Testable: each field maps to something you can inspect in the result.
What you need before you start
Prepare:
- the highest-resolution version of the reference image you can legally use;
- a vision-capable model that can inspect an image;
- an image generator or editor;
- a place to save the analysis JSON, generation settings, and outputs;
- one clear statement of what must stay the same and what may change.
Crop away browser chrome, comments, watermarks that are not part of the intended design, and unrelated borders. Keep the original aspect ratio unless changing it is part of the task.
Current official documentation from OpenAI and Google shows that modern multimodal systems can analyze image inputs. Their generation tools can also use reference images in editing or image-to-image workflows. The exact request format varies by provider, so keep the method below model-agnostic.
Step 1: Extract observation JSON from the reference
Send the image together with this prompt. The wording deliberately blocks invisible details and forces null for unknowns.
You are a visual evidence analyst. Analyze the attached reference image and
return one valid JSON object only.
Rules:
1. Describe only evidence visible in the image.
2. Do not infer hidden causes, the original prompt, exact camera equipment,
exact dates, invisible light technology, or unreadable text.
3. If a value cannot be determined from the image, use null.
4. Keep different visual dimensions in separate fields.
5. Transcribe clearly legible text exactly. Do not guess missing letters.
6. Describe spatial relationships explicitly: left/right, above/below,
foreground/background, centered/off-center, relative size, and overlap.
7. Before returning the JSON, check for contradictions between fields.
8. Use complete, concrete sentences inside values; avoid vague adjective lists.
Return this structure:
{
"image_type": null,
"subject": null,
"action": null,
"location": null,
"composition": {
"orientation_and_aspect_ratio": null,
"framing_and_crop": null,
"subject_placement": null,
"spatial_relationships": null,
"negative_space": null
},
"lighting": {
"visible_direction": null,
"apparent_temperature": null,
"softness_and_contrast": null,
"highlights_reflections_shadows": null,
"unknown_causes": null
},
"color_palette": null,
"materials_and_textures": null,
"camera_and_focus": {
"viewpoint": null,
"perspective": null,
"depth_of_field": null,
"sharp_and_blurred_regions": null
},
"text_and_typography": {
"verbatim_text": [],
"placement_and_alignment": null,
"size_hierarchy": null,
"visible_lettering_style": null,
"unreadable_text": null
},
"style": null,
"mood_and_vibe": null,
"uncertainties": [],
"exclusions": null
}
Validate the JSON before generating
A pretty answer is not enough. Check four things:
- It parses as JSON: no comments, trailing commas, or Markdown around it.
- Required fields exist.
- Unknowns are
nullor listed underuncertainties, not silently invented. - Fields do not contradict one another.
Two examples from the user-reported workflow that inspired this guide show why the last two checks matter. In a refrigerator photo, the visible image supported cool light inside the refrigerator but not the claim that the lamp was specifically LED. In a website screenshot, the analysis correctly found a left-aligned headline but also called the hero centered. One is an unsupported inference; the other is an internal contradiction. Both should fail validation before generation.
Step 2: Turn the observation into an editing plan
Do not immediately rewrite every field. First classify the information into three states:
| State | Meaning | Example |
|---|---|---|
locked | Must be preserved | subject position, visual hierarchy, aspect ratio |
editable | Intentionally changed | product color, setting, headline |
unknown | Not supported by the image | exact lens, hidden light type, font family |
Create a small control object beside the observation JSON:
{
"locked": [
"single main subject in the lower-left third",
"large empty area on the right",
"soft side light and low overall contrast",
"headline above the supporting line"
],
"editable": {
"subject": "replace the ceramic mug with a clear glass bottle",
"accent_color": "change muted red to cobalt blue",
"verbatim_text": ["NORTH", "STILL WATER"]
},
"unknown": [
"camera model",
"exact focal length",
"exact font family",
"physical type of the off-frame light"
],
"hard_constraints": [
"no extra text",
"do not change the camera viewpoint",
"do not add objects outside the reference layout"
]
}
This is more reliable than burying every instruction in one paragraph. It tells the generator what to preserve, what to alter, and what not to pretend to know.
Resolve conflicts before sending the prompt
Run a plain-language consistency pass:
- Can the subject be both centered and left-aligned? If not, choose one.
- Does the crop description match the stated aspect ratio?
- Does “soft diffuse light” conflict with “hard-edged shadows”?
- Is text listed as both absent and required?
- Is an object described in two incompatible positions?
A generator cannot reliably fix a contradictory specification for you. It will choose one interpretation, blend them, or ignore both.
Step 3: Generate a controlled first batch
Use the observation JSON plus the editing plan. When the generator supports a reference image, attach it as well. Official image-generation guides from OpenAI and Google both document image-input or iterative editing workflows, but capability and syntax differ by model.
A practical handoff prompt is:
Create a new image using the attached reference and the JSON specification.
Priority order:
1. Obey hard constraints.
2. Preserve every item in "locked".
3. Apply only the requested values in "editable".
4. Treat "unknown" as unknown; do not invent technical details.
5. Match composition and spatial relationships before fine texture.
6. Render quoted text exactly once, with no extra words.
7. Do not copy watermarks, signatures, or protected logos.
Return one image. Do not explain the prompt.
If the tool ignores raw JSON, convert it to labeled prose without changing the meaning:
SUBJECT:
COMPOSITION:
LIGHTING:
COLOR:
MATERIALS:
TEXT:
STYLE:
PRESERVE:
CHANGE:
DO NOT ADD:
JSON is an editing format, not a universal image-model protocol. Some tools follow natural-language sections better.
For the first batch:
- Match the reference aspect ratio.
- Keep one seed when the tool exposes it, but do not assume seed portability across models.
- Generate three or four candidates, not one.
- Save the exact JSON, model/version, settings, reference file, and output.
- Give each run an ID such as
v01-a,v01-b, andv01-c.
The success signal is not “the image looks good.” It is: at least one candidate passes every hard gate and has no major failure in the highest-priority visual fields.
Step 4: Score the result against the reference
Review each output at the same size as the reference. Use 0, 1, or 2 for every row:
0= wrong or missing;1= partly correct;2= acceptably matched for the intended use.
| Review field | What to compare |
|---|---|
| Subject | count, identity-defining features, silhouette, relative size |
| Composition | placement, crop, balance, negative space, overlap |
| Spatial relations | what is left/right, above/below, in front/behind |
| Lighting | visible direction, warmth, softness, contrast, shadow pattern |
| Color | dominant/supporting/accent colors and their locations |
| Materials | gloss, transparency, fabric, grain, fur, metal, paper |
| Text | exact wording, count, spelling, alignment, hierarchy |
| Style and mood | medium, finish, visual treatment, atmosphere |
Use these hard gates:
- No unsupported claim in the prompt. A guessed “LED,” exact lens, or exact font is not a harmless detail; it can pull the generator away from visible evidence.
- No contradiction in the prompt.
- Required text is correct. If exact wording is essential and still wrong, the image fails.
- Critical spatial relationships are preserved.
- No unrequested logo, watermark, object, or text.
A total score can help sort candidates, but hard gates take precedence. A beautiful image with the wrong headline or reversed layout is still a failed reconstruction for that use case.
Review from coarse to fine
Check in this order:
- canvas and composition;
- subject count, placement, and scale;
- lighting and major color blocks;
- material and texture;
- typography and small details.
There is little value in perfecting fur texture when the subject occupies the wrong half of the frame.
Step 5: Revise one field group at a time
Choose the best candidate as the new baseline. Then edit only the field group connected to the largest failure.
| Failure | Change | Explicitly preserve |
|---|---|---|
| Subject is too large | composition.subject_placement and scale | viewpoint, light, palette |
| Layout drifted | spatial relations and negative space | subject appearance, materials |
| Scene is too warm | apparent temperature and color palette | geometry and text |
| Product looks plastic | materials and reflections | shape, placement, label |
| Text is misspelled | verbatim text, count, placement | all non-text pixels if editing is supported |
| Style is right but identity drifted | subject features or reference strength | composition and background |
Use language such as: “Change only the headline text. Keep the crop, camera viewpoint, object geometry, lighting, colors, and all other text unchanged.”
Official prompting guidance also recommends separating changes from preservation constraints and refining one thing at a time. That makes each iteration diagnosable instead of turning it into a new random generation.
Common failures and how to recover
The analysis model returns invalid JSON
Ask it to repair the existing object, not reanalyze the image:
Repair the following response into valid JSON. Preserve all supported content.
Do not add new visual claims. Return JSON only.
If your tool supports a JSON schema or structured output, use it. Schema validation prevents syntax errors, but it does not prove that the visual claims are true.
The generator ignores the JSON
Flatten the same data into labeled prose. Put PRESERVE, CHANGE, and DO NOT ADD near the end. Remove low-priority detail so the critical constraints are not lost in a long prompt.
Composition keeps drifting
Reduce descriptive style language and add explicit relationships:
- “subject center at approximately 30% of canvas width”;
- “headline starts on the same left guide as the image”;
- “product occupies the bottom third”;
- “right half remains mostly empty.”
Coordinates are guides, not guarantees. Verify the actual result.
Exact text is wrong
Quote the text, specify how many times it should appear, and forbid extra text. Generate the text separately first when the provider recommends that workflow. If correctness is non-negotiable, plan to typeset the final wording in a design tool rather than accepting a near miss. OpenAI’s current documentation explicitly notes that text placement and clarity can still fail; Google’s documentation also advises a text-first workflow for image text.
The same character or product changes between iterations
Use an edit workflow with the best prior image as input instead of generating from scratch. Repeat a short identity block, preserve geometry and labels, and judge multiple runs. Current provider documentation warns that recurring characters, brand elements, and precise layout can still vary.
A request fails before an image is returned
Record the error and request ID. Fix authentication, quota, unsupported input, or moderation problems instead of blindly retrying. Retry only transient rate-limit or server failures with backoff. For a user-correctable prompt or image error, change the request first.
A compact acceptance record
Save this beside every selected output:
{
"run_id": "v03-b",
"reference_file": "reference.png",
"analysis_json": "reference.v01.json",
"generator_and_version": "record-the-actual-value",
"settings": {
"aspect_ratio": "record-the-actual-value",
"quality": "record-the-actual-value",
"seed": null
},
"scores": {
"subject": 2,
"composition": 2,
"spatial_relations": 2,
"lighting": 1,
"color": 2,
"materials": 1,
"text": 2,
"style_and_mood": 2
},
"hard_gates": {
"unsupported_claims": false,
"contradictions": false,
"required_text_wrong": false,
"critical_layout_wrong": false,
"unrequested_elements": false
},
"decision": "accept",
"next_change": null
}
This record turns “I like version B” into an auditable decision. It also tells you whether switching models improved the result or merely changed its style.
Frequently asked questions
Can an image-to-JSON tool recover the exact original prompt?
No. It can produce a useful description of visible evidence. The original prompt may contain discarded instructions, hidden references, seeds, model settings, edits, or post-processing that cannot be recovered from the final pixels.
Should the JSON include camera and lighting details?
Include visible effects: viewpoint, perspective, focus, shadow softness, apparent light direction, and color temperature. Do not assert exact equipment or invisible technology unless that information comes from a separate verified source.
Does every image generator accept JSON?
No. Some accept structured input well; others respond better to labeled prose. Keep JSON as your source of truth, then render it into the format your target tool follows best.
How many iterations should I run?
Start with three or four candidates, select the strongest baseline, and then revise one field group per round. Stop when all hard gates pass and further changes no longer improve the intended use.
Is a high similarity score enough?
No. Similarity can hide decisive errors such as wrong text, reversed layout, an extra object, or a fabricated logo. Use field-level checks and hard gates.
Sources and evidence level
- Vox’s X thread: a user-reported practice that motivated the field-by-field approach; it is not an independent benchmark.
- OpenAI Images and vision, image generation, and image prompting: official documentation for image inputs, iterative edits, prompt structure, evaluation, and known limitations.
- Google image understanding and image generation: official documentation for image analysis, structured JSON responses, reference-image generation, iteration, and current limitations.
The method is intentionally provider-independent: keep observable evidence in one structured record, keep changes explicit, and judge the output against the reference rather than against your memory of the prompt.