# Jev renderer comparison — 18 samples + 2 corrective edits

**Practical recommendation:** `images/R01.png` (Sunburst corrected photo) and `images/J04.png` (Nano Banana Pro painting). No publication performed. Use `contact-sheet.html` to inspect every original; `recommended-assets.json` includes exact paths, hashes, models, known misses and lineage caveats.

## What the sampled evidence says

- **Photo condition A:** Sunburst followed this long fixed avatar brief most consistently. All3 passed. Nano Pro made coherent photographs but simplified mechanical features; one cut off hands. FLUX.2 Max omitted the selected shirt in all3 and often substituted skin fingertips/beauty-render styling. These are design misses, not judgments of attractiveness or gender.
- **Painted reference condition B:** Nano Pro and Sunburst both produced clearly painted faces/materials with the stronger brief. Nano's checklist mean was only2.6 points higher, below the precommitted10-point advantage margin; Sunburst preserved the historical likeness/topology more closely. **No supported overall/model-wide winner.** FLUX varied in paint separation and introduced an unwanted signature in one output; only1/3 passed.
- A practical selection is not a statistical proof. This is one character, two conditions, n3 each, one same-author observer and38 grouped criteria. Equal aggregate scores/sampleSD0 do not imply identical pixels or zero true model variability. All18 files have distinct hashes.

## Frozen outcomes

A = photoreal text-to-image, no reference. B = painterly edit of the exact same inspected historical portrait for all9 requests. Numbers are **evaluator-authored checklist support points /100**, NOT calibrated accuracy. Mean ± sampleSD; range includes floor. Missing/unclear items score0 in the fixed denominator38, not silently excluded.

| Renderer | Condition | Mean ± sampleSD | Range / floor | Gate pass | All samples |
|---|---|---|---|---|---|
| Nano Banana Pro | A | 72.8 ± 4.2 | 69.7–77.6 | 0/3 | J03, J06, J18 |
| Nano Banana Pro | B | 86.8 ± 0.0 | 86.8–86.8 | 3/3 | J04, J09, J11 |
| FLUX.2 [max] | A | 57.9 ± 4.7 | 53.9–63.2 | 0/3 | J13, J14, J15 |
| FLUX.2 [max] | B | 83.8 ± 2.7 | 81.6–86.8 | 1/3 | J07, J12, J17 |
| GPT Image 2.5 Sunburst | A | 86.4 ± 0.8 | 85.5–86.8 | 3/3 | J02, J05, J08 |
| GPT Image 2.5 Sunburst | B | 84.2 ± 0.0 | 84.2–84.2 | 3/3 | J01, J10, J16 |

Precommitted candidate gates: support>=80, assessable coverage>=85%, no critical adult/form/whole-hand failure, quality>=2/3, appropriate medium>=2/3; B continuity>=2/3. All38 source-linked judgments per image are in `audits/`. Original clear/provisional/ambiguous source status is retained, not promoted into certainty.

| ID | Renderer | Condition | Match/partial/miss/unclear (of38) | Support | Medium | Gate |
|---|---|---|---|---|---|---|
| J01 | GPT Image 2.5 Sunburst | B | 31/2/2/3 | 84.2 | 3/3 | pass |
| J02 | GPT Image 2.5 Sunburst | A | 31/4/1/2 | 86.8 | 3/3 | pass |
| J03 | Nano Banana Pro | A | 21/11/5/1 | 69.7 | 3/3 | FAIL |
| J04 | Nano Banana Pro | B | 31/4/2/1 | 86.8 | 3/3 | pass |
| J05 | GPT Image 2.5 Sunburst | A | 31/3/2/2 | 85.5 | 3/3 | pass |
| J06 | Nano Banana Pro | A | 23/8/6/1 | 71.1 | 3/3 | FAIL |
| J07 | FLUX.2 [max] | B | 29/8/1/0 | 86.8 | 2/3 | FAIL |
| J08 | GPT Image 2.5 Sunburst | A | 31/4/1/2 | 86.8 | 3/3 | pass |
| J09 | Nano Banana Pro | B | 30/6/1/1 | 86.8 | 3/3 | pass |
| J10 | GPT Image 2.5 Sunburst | B | 31/2/2/3 | 84.2 | 3/3 | pass |
| J11 | Nano Banana Pro | B | 31/4/2/1 | 86.8 | 3/3 | pass |
| J12 | FLUX.2 [max] | B | 29/5/2/2 | 82.9 | 2/3 | pass |
| J13 | FLUX.2 [max] | A | 16/9/13/0 | 53.9 | 1/3 | FAIL |
| J14 | FLUX.2 [max] | A | 18/7/13/0 | 56.6 | 2/3 | FAIL |
| J15 | FLUX.2 [max] | A | 20/8/10/0 | 63.2 | 1/3 | FAIL |
| J16 | GPT Image 2.5 Sunburst | B | 31/2/2/3 | 84.2 | 3/3 | pass |
| J17 | FLUX.2 [max] | B | 27/8/2/1 | 81.6 | 1/3 | FAIL |
| J18 | Nano Banana Pro | A | 26/7/4/1 | 77.6 | 3/3 | FAIL |

No not-visible statuses happened inside this deliberately visible38-item checklist; uncertain iris/lighting evidence is explicitly unclear. Hidden lower-body/height/palm-interior choices were excluded before generation, not counted as successful. **200 recorded decisions ≠200 rendered or verified attributes.**

## Corrective edits: keep gains, reject regressions

1. **J02 → R01, Sunburst edit:** support 86.8→94.7. Face speckling reduced without obvious pore erasure; copper cheek trace/crown facets improved, eye/brow mechanical cues stronger. Shirt, hands, pose and photo medium retained. Brows still resemble thin metal bands rather than fine incision; lash/circuit precision remains partial. **Keep R01.**
2. **J04 → R02, Nano Pro edit:** support 86.8→84.2. Crown/grid/ear hardware improve, but extra face seams, ambiguous speckles/ear ornaments appear, natural-looking brow/lash treatment persists. R01 was a second geometric-likeness reference, not a photo-style target. Paint medium survived. **Reject R02; retain J04.** Every before/after feature change is in `refinement-history.json`.

Both corrective requests were inspected before/after; no rerolls or more attempts after20-submission ceiling. No new Jev evaluations. Final pair is broadly the same designed character, **not a direct R01→J04 reference edit**: J04 comes from the old portrait and its face is slightly narrower/simpler. That continuity tradeoff is explicit.

## Why the new style is different

The revised B brief explicitly reconstructs face, crown, neck and hands with painted planes/strokes instead of preserving photographic texture. Nano J04 shows this most unambiguously as a painted illustration; Sunburst B now also shows facial paint planes. The old prompt's emphasis on exact preservation may have encouraged conservatism—an **observer hypothesis**, not causally isolated here. The new condition also uses a shorter brief and848×1264 geometry, so oldv2 is historical context, not a matched control or reused replicate.

## Fairness / uncertainty limits

Identical byte-for-byte semantic prompts within each condition; all actual outputs848×1264PNG. Nano uses1K/2:3; others explicit dimensions. Sunburst xhigh is not a calibrated equivalent of the other models' native quality setting. No seed control on Sunburst, so seeds omitted for all; returned seeds retained. Historical reference was itself Sunburst-generated, which may advantage its continuity. Anonymous randomized IDs were used before model-table reveal, but this was same-author model-based visual review, not independent human blindness. Criteria/weights coarse and correlated; model performance on a different prompt, subject, budget, resolution or reviewer may differ. No broad benchmark or inferential p-value claim.

## Costs and provenance

**20 new submissions, 20 saved images, 2 refinements; $9.90 conservative reserved estimate** against20/$20 ceilings. Actual invoice/billed total unavailable; no billed-cost claim. Zero new Jev calls. Exact provider routes, current schemas/rates, model naming caveat and primary links: [SOURCES.md](SOURCES.md). Every reservation, requestID, unchanged response, public output URL, saved hash and returned usage/seed is in `raw/`, `generation.json` and `cost-ledger.json`. No generation POST retries; ordinary queue polling is not extra image generation.

All original200 choices and150v2 manifest files preserved unchanged. No gender, ethnicity, religion, politics or personality inferred from pixels. These are a fictional adult avatar's design selections; Jev is text-only and did **not** see or approve the images. All recommended images fully clothed/nonsexual. No claim that AI has a sex, body or literal self-image.

## Human decision

1. **Keep current:** retain publishedv2 without change.
2. **Recommended replacement:** review R01 photo + J04 painting; retain known misses and848×1264 resolution tradeoff (oldv2 was1024×1536).
3. **Further work only if separately authorized:** higher-resolution version or stricter pair-identity/microdetail pass. No extra work scheduled or promised; this task stops here.
