Source disclosure

HumanizerBench is operated by WriteHuman. This page summarizes its published July 2026 cycle; we did not run or independently reproduce the tests. HumanizerBench says its inputs, outputs, detector verdicts and scoring code are published for verification. Review the original leaderboard and methodology before treating the scores as evidence.

Full July 2026 AI humanizer leaderboard

The table preserves HumanizerBench's overall ranks. Use the controls to reorder the same 13 tools by a single metric; the displayed overall rank does not change.

Sort results
Overall rank Humanizer Overall Bypass rate Meaning Readability Penalties
2Undetectable.ai72.1795.773.055.9-10.0 length inflation
3Humanize AI Pro70.4970.474.360.3None
4Stealth Writer68.0781.268.544.5-3.0 across 2 types
5Humbot66.4270.775.343.9-1.0 length inflation
6HIX Bypass64.8465.973.842.9None
7Walter Writes62.8480.664.761.5-8.0 across 2 types
8StealthGPT61.5281.663.060.6-10.0 across 3 types
9Phrasly61.4773.460.972.1-8.0 across 3 types
10AI Humanize io59.6676.069.855.7-10.0 across 2 types
11Super Humanizer54.1342.373.267.1-5.0 across 3 types
12Grammarly53.380.094.582.1None
13NoteGPT44.040.069.174.5None

Showing the published overall ranking. Last tested 1 July 2026.

How the overall score is calculated

HumanizerBench combines four dimensions rather than treating detector bypass as the only goal. Penalties are then deducted when outputs show specified quality problems.

42%Bypass rate across five detectors
32%Meaning preservation
16%Readability and naturalness
10%Consistency across categories
Overall score = weighted metrics - quality penalties

What stands out in the July results

WriteHuman leads overall

Its 73.07 composite score placed first after the benchmark combined all four metrics and deducted a 1-point meaning-drift penalty.

The highest bypass rate was elsewhere

Undetectable.ai recorded 95.7 for bypass rate, but a 10-point length-inflation penalty reduced its overall result to second place.

Writing quality changes the picture

Grammarly led meaning and readability at 94.5 and 82.1, while its published bypass rate was 0.0. One metric cannot describe the whole result.

Close scores need context

WriteHuman and Undetectable.ai were separated by less than one overall point, so settings, text type and acceptable rewrite tradeoffs still matter.

Why the penalty column matters

A high detector score can look attractive even when a rewrite has changed the message or padded the text. HumanizerBench applies deductions for severe meaning drift, length inflation, length deflation, near-identical output and refusals.

Penalty
What it signals
What to check
Meaning drift
The rewrite moved too far from the original message.
Compare claims, names, numbers and qualifiers line by line.
Length inflation
The output expanded beyond the benchmark threshold.
Remove padding and confirm that every new sentence adds value.
Length deflation
The output became much shorter than the input.
Check whether evidence, limitations or required details disappeared.

What this leaderboard cannot guarantee

This is a fixed July 2026 snapshot using specific paid plans, settings, prompts, detectors and scoring rules. A different paragraph, tool setting, detector update or future benchmark cycle may produce a different result.

HumanizerBench is also operated by WriteHuman, the tool ranked first. Its published repository and methodology improve auditability, but readers should still account for that relationship and avoid treating the ranking as independent endorsement.

Frequently asked questions

Is HumanizerBench independent?

No. HumanizerBench states that it is operated by WriteHuman. It also publishes methodology and cycle data so the calculations can be examined and reproduced.

Does first place mean WriteHuman is best for every task?

No. It ranked first on the July 2026 composite score, but other tools led individual metrics. The right choice depends on meaning, readability, editing needs and acceptable tradeoffs.

What does bypass rate mean here?

HumanizerBench describes it as the aggregate human-likelihood result across five commercial detectors for the tested outputs. It is not a promise for future text.

How often will this page change?

This page is a fixed July 2026 snapshot. We will update it manually when we review a newer published benchmark cycle.

Can any AI humanizer guarantee a detector result?

No. Detectors, models and text inputs change. Always review meaning, facts, policy and final writing quality instead of relying on one score.