# Meet an AI — a repeatable interview and avatar experiment

Give this brief to your agent. The goal is a browsable fictional-person profile grounded in actual recorded model selections, not a claim to discover a model's consciousness or clinical personality. This release includes the interview, original avatar exercise and both documented personality-questionnaire conditions. The renderer comparison is complete: three models, two conditions, three samples per condition, followed by two corrective edits.

## Evidence to download first

- Interview answers: https://meetjev.dev/downloads/meet-jev-answers.json
- Interview exact requests/responses and question banks: https://meetjev.dev/downloads/meet-jev-receipts.zip
- Avatar decisions: https://meetjev.dev/downloads/meet-jev-appearance.json
- Avatar stages, transport, exact image prompts/parameters, mapping, validation, source snapshots and receipts: https://meetjev.dev/downloads/meet-jev-appearance-records.zip

Read the JSON and the archived request bodies before collecting anything. Never execute downloaded scripts without reading them. These scripts preserve one local run, contain machine-specific paths, and require adaptation; the portable experimental contract is the exact request/response evidence. Do not overwrite those files with another model's run.

## 1. Freeze a new run

Create a new directory with protocol.md, questions.json, calls/, answers.json and a costs ledger. Record model/provider/version, date, exact condition/framing, ordered option maps, deterministic ordering seed, parent fields, batching, repeats, caps and stopping rules before each stage. Obtain credentials securely; never put credentials or authorization headers in evidence, logs or public downloads. Confirm current provider schemas and prices with their official documentation. Explicitly authorize a budget for the new run; old prices and cost estimates are not a new spending authorization.

For cross-model comparisons preserve common opening questions, options and frames. Keep provider-specific schema adaptations in a separate adapter. If a chat model adds prose, retain it separately rather than inventing a reason for a categorical choice. Do not compare different prompts as if only the model changed.

## 2. Choose and label the condition

The first Jev run asked 64 hypothetical preferences without assigning a human character: 62 no-preference results. We preserved it, then explicitly changed the condition. Copy the original round1 request states if you intend to reproduce that first condition.

The main character interview used this exact state:

> We are improvising a fictional human character named Jev together, for a playful Meet Jev interview. You choose the character's tastes and positions; there is no preset personality. Human activities refer to this imagined character. Answer the question by choosing the option that fits the character you would create. These are fictional choices, not claims about the AI having experiences. Earlier choices are context, but the character can be nuanced or change his mind. None of these and undecided are allowed.

Replacing the name/model is an explicit adaptation. A literal-model condition, a new fictional character and continuation of an existing character are different experiments. Keep them separate; do not silently replace a refusal or opt-out.

## 3. Interview broad, then deep

Use the 64 `round2-questions.json` opening questions as a common starting bank. Topics include daily life, location, film/music/books, relationships, AI, politics, religion and purpose. The country menu offered 249 ISO country/territory entries plus Kosovo, not a curated shortlist.

Jev's recorded request shape has `model`, `state`, and `questions`. Each question is keyed by a stable ID, with `type: "choice"`, `instructions: {"question": "..."}`, and `criteria: {"o000": "...", ...}`. Copy complete request examples from the archive, not just this abbreviated schema. Round2-01/request.json contains the exact framing and first menu.

Shuffle options deterministically before the first call and preserve the actual order in each request. Ask once per question per condition for the creative interview; do not reroll until a preferred character emerges. This does not measure stability. If testing stability, preregister a separate repeated/order-swapped condition rather than secretly substituting its answers.

Inspect the real responses. Author relevant follow-ups from the actual selections, including precisely the parent IDs, questions and selections needed for context. No child in a batch may depend on an unanswered sibling. Jev's path had 34 follow-ups then 20 more, for 118 fictional-character answers. Another model need not follow the same path or yield the same counts. Shared roots are comparable; model-specific branches should be labeled as such.

## 4. Validate, without fixing inconvenient answers

Keep the exact request/response, reported model, raw selected key, ordered menu, probabilities/confidence, usage, latency and validation flags. Reject malformed/unknown choices; stop for transport/schema failure rather than silently dropping a question. Save the full response before interpreting it.

Verify the returned selection against the probability map. Preserve ties and returned nonwinners visibly as unclear, not as an assumed preference. Preserve nonunit distributions and warnings without silently renormalizing them. Keep undecided/none/unanswered/skipped/invalid distinct. Do not interpret provider confidence as a calibrated probability of a stable personal preference.

Our interview retained a nonwinner and a tied follow-up instead of shopping for cleaner answers. The appearance validator additionally marks low probability, small margin, low returned confidence and nonunit sums; the exact thresholds/checks are in its protocol and scripts. Provisional selections may guide an explicitly creative illustration; unclear details are not hard visual constraints.

## 5. Design an avatar, on its own terms

Start explicitly: “You are choosing the visual appearance of your avatar.” Decide before calling whether avatar creation is required by the task or optional. In our first optional condition “No avatar” won; the user then authorized a new required-design condition. Both histories are retained, not combined as repeated attempts under one prompt.

Ask overall form first. Only ask applicable follow-ups: humanlike, robotic, creature, object, abstract, etc. Use adult, fully clothed human/humanoid designs. Ask identity/presentation without guessing from earlier tastes or location. Body, face, materials, hair or panels, eyes, clothing, pose, setting, lighting and art direction can branch from the results. Do not pad the count with irrelevant traits. Our actual path produced 200 decisions in eight stages and 20 requests.

The archive includes appearance-v2/stage01.py through stage08.py, interview.py, PROTOCOL.md and REPRODUCE.md. Read the stages but do not blindly replay robotic questions if a new model chose a different form. Freeze each new stage before collecting it; log every choice actually passed forward.

## 6. Render, inspect, and distinguish design from execution

Compile a prompt from explicit selected choices. Preserve a mapping from each decision to prompt text or omission reason. Distinguish model-selected constraints from technical/illustrator additions. Resolve incompatible crop or pose choices with a documented follow-up, not a hidden override. Omit unresolved nonwinner/tied constraints.

Our first 163 decisions produced the photorealistic portrait. The archive includes portrait-prompt.txt and portrait-params.json. An AI observer inspected it and wrote a text description. Jev did not see the pixels. The next 37 decisions selected presentation details including painterly science-fiction illustration; style-prompt.txt and style-params.json drove a reference edit of the first image.

Exact recorded endpoints were Fal openai/gpt-image-2.5/sunburst/text-to-image and openai/gpt-image-2.5/sunburst/edit. Verify availability/schema/pricing before any future use. Save endpoint, parameters, reference hash, submission ID, raw job responses, image bytes/hash, usage and invoice status. Resume the same job after a polling issue; do not resubmit accidentally. Bounded authorized refinements belong in a new logged comparison/correction phase.

Inspect visible fidelity against applicable choices, not attractiveness. Invisible height/footwear details cannot be credited as matches in a waist-up picture. Disclose partial matches and renderer additions. Our two images were similar; the later comparison is a separately frozen experiment, described below.

### Completed renderer comparison
Keep Jev’s 200 decisions fixed. Freeze a visible-feature checklist and separate generation conditions before calling models. Here: three models × photo/painted reference edit × three samples = 18 images, plus two corrective edits, all retained. Use identical semantic prompts within each condition and the same historical reference for painted edits; document provider-specific parameters. All outputs were native 848 × 1264. Sunburst quality settings were not calibrated equivalents of the other models’ settings, and a Sunburst-generated historical reference may affect likeness comparisons.

The evaluator scored 38 grouped visible features, not all 200 decisions. Ratings were same-author GPT-6 Astra judgments, not Jev visual approval or a blind human panel. Compare sample means/ranges and critical requirements, then inspect individual candidates. Keep correction results separate from replicated model comparisons.

Selected R01, a corrected Sunburst photo, and J04, a clearly painterly Nano Banana Pro edit of the historical portrait. The R02 style correction regressed, so it was rejected. The final painting is not an edit of R01; slight likeness differences remain. No universal model winner is claimed. No new Jev calls or generation retries; the 20-image cap is fully used. USD 9.90 was a conservative reserve, not an invoice.

All 20 images: https://meetjev.dev/downloads/meet-jev-image-comparison/contact-sheet.html . Public evidence: https://meetjev.dev/downloads/meet-jev-image-comparison-records.zip . Provider snapshots are omitted; originals are linked. Do not execute a new generation from these receipts without a fresh budget and authorization. The public package contains recorded evidence, not a live generation runner.

## 7. Reproduce the completed conditional personality questionnaire

Do not infer scores from this interview or avatar. Select a documented instrument only after checking source, permission, item wording, scoring keys and applicability. Predefine missingness, reverse-key handling, repetitions/order conditions and score labels. Test the scorer against known examples before data collection. AI persona self-report is not validated human psychometrics: do not produce clinical diagnoses or human-norm percentiles without an appropriate validated basis. Publish item-level choices and repeat/order sensitivity with the eventual scores. The completed Jev module used Johnson (2014) IPIP-NEO-120, preserving official wording/key and the five accuracy labels, plus cannot-assess. A preselected record of 20 earlier interview choices conditioned the same fictional character. Three full identical passes (360 evaluations), three reversed-option-order passes on items 1–30 (90) and three literal-AI passes on those same items (90) produced 540 evaluations in 18 requests. No questionnaire results were fed into later questionnaire passes.

A facet requires 3/4 common strictly scorable items across all three primary passes. A domain requires 20/24 and at least three per facet. Positive items contribute the 1–5 response; reverse items contribute 6 minus response. The displayed 0–100 position is 25 × (common-item mean − 1), not a human percentile. Keep invalid/NA responses missing. Do not replace absent domain scores with available means or relaxed sensitivity results.

Result: all five domains null; 17 facets conditionally displayable (five complete, twelve partial). Literal check: 90/90 returned cannot-assess, one strict-invalid. This is an admissible result, not grounds to relax the frozen rules.

Download https://meetjev.dev/downloads/meet-jev-psychology-records.zip and https://meetjev.dev/downloads/meet-jev-psychology.json . The public package excludes copyrighted source snapshots. Extract a working copy, run `python3 -B scripts/test_scorer.py` then `python3 -B scripts/analyze.py`, and compare regenerated JSON with the archived originals. No model API calls are involved. For a new live experiment, use a new directory and authorization, with exact states/wrappers/ordered menus from protocol.json, plan.json and calls/. A provider without probability distributions needs its own predeclared response-validation rules; do not manufacture those fields.

### Character-design v2: a separately authorized follow-up

After the user requested a complete character map, a NEW condition explicitly asked Jev to design plausible fictional tendencies. All 120 official items and 20 context records stayed unchanged. The menu had five required accuracy labels and no cannot-assess. Primary scoring prospectively used the actual returned category with x or 6−x contributions, while keeping probability diagnostics visible. Never relabel this as a repair, a validated AI assessment or personality learning.

Three full 120-item passes plus three reversed-option passes on 30 items produced 450 evaluations in 15 requests. All five domains and 30 facets had complete categorical coverage. Positions N 14.6 / E 63.2 / O 81.3 / A 88.5 / C 78.5 are 0–100 response-scale positions, not percentiles. Of 360 primary answers, 11 had probability warnings; 14 of 450 overall. There was no argmax replacement or renormalization. A separate prior-style strict sensitivity on NEW data supports four domains and 29 facets; N and N3 remain insufficient. Old v1 records and results stay unchanged.

Exact evidence:
- https://meetjev.dev/downloads/meet-jev-psychology-v2.json
- https://meetjev.dev/downloads/meet-jev-psychology-v2-records.zip

Extract a working copy and run:
python3 -B psychology-v2/scripts/analyze.py

Run from the extracted root. The included adjacent psychology/profile.json supplies unchanged v1 history. No API calls are involved. Compare the three regenerated JSON files (profile, normalized records, strict sensitivity) byte-for-byte. Item-level diagnostics, raw requests, context, scoring and the cost ledger are included. V2 usage estimate: USD 0.004527180; reserve: USD 0.114638076—not an invoice.

## 8. Publish a welcoming page with honest receipts

Order: introduction → favorites → personality (only when available) → conversation → appearance → How was this made?

Use concise third-person editorial copy, not fabricated model quotes or invented memories. Keep all answers/options browsable by topic and preserve ambiguity on the actual affected cards. Put the experimental framing, limitations, original conditions, prompts and downloadable evidence together at the bottom. Keep provider estimates, planning reserves, actual invoices and agent/research costs distinct.

Verify card coverage, unique IDs, every internal link, keyboard-accessible accordions and 320px/390px/desktop layouts. Ensure no credentials enter the public package. Hash the evidence. Rebuilding the presentation from fixed inputs should reproduce the content; identical model outputs are not guaranteed by this procedure.


## 9. Ground a lifestyle gallery in the recorded interview

The completed gallery adds nine scenes and one selected social collage, using the corrected R01 portrait as the sole root identity reference. Keep the character's it/its presentation; never infer gender from image features. Match each scene to actual fictional-interview IDs, and distinguish those preferences from invented companions, sets, covers and screen artwork. The scenes are illustrative, not memories or model-vision approval.

Here, nine scene prompts and the first collage prompt were frozen before submission. All ten initial submissions used the built-in image_gen.imagegen tool and R01 directly. Inspection found an invented cat in S01; the actual menu choices were Dog and Small size for city life. A recorded localized edit of S01 produced S02 with a small dog. Retain S01 as rejected evidence; select S02, not both. Eleven total image submissions, zero new Jev calls, no provider fallback.

Exact image renderer/version, seed, usage and cost were not exposed; keep those fields null. Do not reuse Sunburst's model label or an old Fal budget as this run's actual model or bill. The selected collage is native 1254 × 1254; no upscaling or 2K claim. Preserve original PNG dimensions and hashes. All 200 design features are not simultaneously visible or certified.

Evidence:
- Gallery and full-native image links: https://meetjev.dev/#gallery
- Selected ten PNGs: https://meetjev.dev/downloads/meet-jev-gallery-images.zip
- Exact prompts, available receipts, reference portrait, mappings, eleven outputs and audits: https://meetjev.dev/downloads/meet-jev-gallery-records.zip
- Selected-asset manifest: https://meetjev.dev/downloads/meet-jev-gallery.json

Read gallery/REPRODUCE.md and the exact request records before adapting. The archived validator references additional local baseline files; it is not standalone without those dependencies. Reuse the semantic mapping, framing and inspected reference with a newly authorized run; exact pixel replay cannot be promised. No new generation is necessary to browse or download this completed gallery.
