hold your voice vs gpt-5.6 for writing
a blind benchmark of regular gpt-5.6 drafts and profile-aware hyv edits across 16 matched writing tasks.
a strong writing model can give you a clean first draft. that is a different job from holding onto the way a specific person notices, argues, and lands a point.
we wanted a test built around fixed inputs and blind labels. so we froze the tasks and source facts, generated regular drafts with gpt-5.6, applied the hyv treatment only after that, hid the labels, and asked two independent reviewers to choose.
blind reviewer decisions preferred the hyv-edited draft.
paired comparisons had exact agreement between both reviewers.
mean improvement in the composite voice score on a five-point scale.
what did the gpt-5.6 writing benchmark test?
the test measured whether profile-aware editing improves a regular gpt-5.6 draft when the content assignment stays the same. both conditions used the same prompt, source facts, and blinded candidate labels.
| part | what we used | why it matters |
|---|---|---|
| writing tasks | 8 fixed tasks with a brief, audience, format, and source facts. | each pair solves the same writing problem. |
| models | gpt-5.6 sol and gpt-5.6 terra. | the result covers two generation models. |
| draft pairs | 16 regular-prompt and hyv-edited pairs, 32 drafts total. | every comparison uses a direct baseline. |
| blind review | 2 independent gpt reviewers, 64 blinded candidate reviews. | reviewers saw randomized candidate labels. |
| editing boundary | 30 changed lines total; unflagged lines were preserved. | the treatment targeted voice drift instead of rewriting the assignment. |
how was the comparison kept fair?
both conditions began with the same task packet. the regular condition stopped at the gpt-5.6 draft. the hyv condition used that exact draft as input for a profile-aware, line-targeted edit. source facts and required constraints stayed in the task packet.
| step | regular gpt-5.6 draft | hyv-edited gpt-5.6 draft |
|---|---|---|
| 1. write | gpt-5.6 received the brief and source facts. | the same gpt-5.6 draft and task packet were used. |
| 2. assess | no profile-aware pass. | hyv scanned for generic ai patterns and voice drift. |
| 3. edit | the original model draft remained the candidate. | only flagged lines received a targeted rewrite. |
| 4. review | candidate labels were hidden before the blind reviewer pass. | |
what were the benchmark results?
hyv was the preferred candidate in 29 of 32 blind decisions. both gpt-5.6 sol and gpt-5.6 terra showed a positive voice-score delta after the profile-aware edit.
| measure | regular prompt | hyv edit | change |
|---|---|---|---|
| blind reviewer preference | 3 decisions | 29 decisions | +26 decisions |
| mean voice composite | 2.97 / 5 | 3.31 / 5 | +0.34 |
| mean generic quality | 3.66 / 5 | 3.98 / 5 | +0.33 |
| mean hyv score | 2.50 / 100 | 93.75 / 100 | +91.25 |
| mean detected pattern count | 2.19 | 0.06 | -2.13 |
which voice qualities improved most?
the gain was concentrated in parts of writing that a generic model can smooth out: distinctive phrasing, the shape of the argument, and sentence movement. this is the difference between readable copy and copy a repeat reader can recognize.
| dimension | regular prompt | hyv edit | delta |
|---|---|---|---|
| voice specificity | 1.94 / 5 | 3.13 / 5 | +1.19 |
| signature architecture | 2.47 / 5 | 3.16 / 5 | +0.69 |
| rhythm | 2.84 / 5 | 3.47 / 5 | +0.63 |
| tone boundary | 3.53 / 5 | 3.91 / 5 | +0.38 |
why does a profile-aware edit change gpt-5.6 writing?
a strong prompt tells a model what to make. a voice profile gives the edit pass a reference for what should remain recognizably yours. hyv uses that reference to find repeatable ai patterns and limits its intervention to the lines that triggered the check.
the aim is to retain the useful parts of the model draft while restoring the choices that make a writer identifiable. our guide to making ai writing sound like you explains the role of real samples; the ai writing patterns guide shows the patterns this pass is built to catch.
what does this benchmark measure?
this is an early product benchmark, with eight tasks and model-based blind review. it gives us a controlled answer to one narrow question: with the same gpt-5.6 assignment, does the hyv edit make the result more voice-specific to blinded reviewers? in this pilot, yes.
the next study should add human authors and editors, more writing formats, and a pre-registered scoring rubric. the repeatable part is already clear: freeze the inputs, blind the candidates, separate voice from factual review, and publish the method with the result.
how can you run the same ai writing test?
start with one writing task your team actually ships. the goal is a decision you can trust.
- freeze the brief, target reader, and source facts.
- generate a regular model draft and a profile-aware edit from that same draft.
- remove labels and randomize the candidate order.
- score voice match, factual preservation, and editor preference separately.
- keep the source facts beside every review so a smoother draft never wins by inventing.
you can begin with the free brand voice analyzer, then use a brand-voice measurement rubric to keep editorial decisions separate from generic model quality.
does hold your voice replace a strong gpt-5.6 prompt?
no. a strong prompt remains the starting point. hyv is the profile-aware layer after the first draft: it checks whether the model is drifting toward generic patterns and supplies a bounded edit where the draft loses the writer's identifiable choices.
get started for $1 — create your account and scan your first draft in minutes.
get started for $1 →






