First double-crown holder for

Novelty & Utility

NoveltyBench · Carnegie Mellon University

Ranked first for both novelty and utility by the Carnegie Mellon NoveltyBench, which evaluates LLMs for humanlike diversity.

The Illusion of the Creative Prompt

The Limits of Statistical Familiarity

Frontier AI models (like the ones powering ChatGPT, Gemini and Claude) are designed to play it safe. They naturally gravitate toward the most familiar, statistically correct answers.

Most of the time, that is exactly what you want. But when you are hunting for a breakthrough, this becomes a roadblock. Human brainstorming thrives on exploratory, divergent thinking. AI, however, thinks convergently. It leans heavily on prior art, rehashing existing ideas and churning out the same tropes within a narrow, predictable bandwidth over and over again.

We set out to fix that.

Why Prompting for Creativity Fails

You cannot just tell a model to "be more creative". Because its core imperative is to be correct, the AI will immediately steer you back to the most statistically average, and therefore uninteresting, response.

Imagine your car refuses to start. Ask an AI what to check, and it will give you the standard, highly probable checklist. It will not suggest that your immobiliser might have forgotten its key. Yet, in your specific circumstance, that rare, out-of-the-box check might be the exact fix you need.

Defining the AI Utility Tax

People often assume out-of-the-box thinking is just for creative ad campaigns, but that same divergent leverage is critical for cracking hard science, strategy, finance and legal problems.

The catch? The industry had accepted a fundamental trade-off. As shown by CMU's NoveltyBench paper and Dr Jiang's Artificial Hivemind research, making an AI more creative directly degraded its practical usefulness. You paid a "utility tax": the more novel the model became, the less accurate and applicable its output was.

Breaking Out

Novamine is the first AI globally to eliminate this trade-off, taking both the Novelty (distinctness) and Utility crowns simultaneously. This proves that:

  • Novelty does not need to be artificially limited.
  • Distinctness does not have to come at the cost of practical accuracy.
  • Instead of relying on older models with low utility, this can be applied directly to today's most powerful frontier AIs with zero loss of accuracy.

This creates the first scientifically proven doorway to safely explore divergent, alternative approaches to your most intractable challenges. It gives you access to solutions your competitors simply cannot see.

The Record We Hold

Before Novamine, achieving a high novelty score required fundamentally altering a model's wiring, which ruined its practical usefulness. Ours runs on unmodified, full-strength commercial AI. Carnegie Mellon reproduced our curated-set numbers through the public endpoint in September 2026, and the entry was merged onto the leaderboard on 9 September 2026.

THEIRS · Best Other Entry on the Board
6.68 Distinct   5.32 utility
Claude Fable 5.1 driven by the benchmark authors' own regeneration prompting (10 September 2026). The best raw model scores 2.43 and 2.90.
OURS · Novamini Brainstorm, the Official Entry
9.30 Distinct   6.65 utility
+39% and +25% over the best other entry, 3.8x and 2.3x the best raw model, on the same 1,100 prompts and the same judges (v1.1, September 2026), from one entry. The double crown.

What This Page Lets You Test Drive

Novamine, the full system, is a 24-hour-plus engagement: a pipeline of generation and agentic adversarial verification designed as a human-in-the-loop deep R&D acceleration.

Novamini has the same core mechanisms distilled into a run of minutes rather than hours, designed for purely creative exploration, or for the initial scientific or strategic brainstorming that leads your team to bigger, more rigorous ideas.

The Science five minutes, in simple terms

Measured, Not Asserted

Novamine sits between you and the model you already use and draws out answers that are measurably new and measurably useful. Nothing is retrained, and we only claim what we can prove with hard scores: Carnegie Mellon's public benchmark ranks it first on both counts.

First & Only AI to Hold the Double Crown on Carnegie Mellon's NoveltyBench

NoveltyBench is the public benchmark built by researchers at Carnegie Mellon University's Language Technologies Institute, one of the world's leading centres for the science of language models, to settle the question that matters here. Ask an AI the same question ten times: how many of the ten answers are genuinely different ideas (the benchmark calls this Distinct), and how useful are they, added up (utility)? Its judges are trained models calibrated on thousands of human judgements, its scoring code is public, and its leaderboard admits only systems the Carnegie Mellon team can reproduce themselves. Holding the top score on both at once is the double crown, and until now no entry had done it.

What the test involved: 1,100 briefs, each answered ten times, meaning 11,000 scored answers per entry, using independent AI judges to ensure no model marks its own homework. The complete output set and scoring records are preserved as an evidence pack, available on request. This is not a demo video; it is the most expensive kind of evidence we could buy.

Independent Eyes: The Hivemind Result

In 2025, Dr Liwei Jiang and colleagues published the Artificial Hivemind, an award-winning paper that put a hard number on the repetition described at the top of this page: on open-ended questions, 79% of prompts drew near-identical answers across fifty generations. Dr Jiang is regarded as one of the most prominent scientists in the world on the behaviour of large language models.

We replicated her study with her own instrument, then applied Novamine: same models, same questions, nothing tuned per prompt. The repetitive answers practically vanished. On GPT-4o, prompts above the paper's similarity line fell from 96% to 1%, and the same unmodified system repeated the effect on a second provider's model.

Repetition, it turns out, is not a law of AI. It yields to Novamine, which mines the model's full depth without retraining it.

Four-arm chart: raw versus orchestrated similarity on GPT-4o and Claude Opus 5
Fig. 1: mean pairwise similarity per prompt (50 generations each, 100 prompts per arm), protocol of Jiang et al. Raw models sit above the 0.8 "hivemind" line on most prompts; the same models under Novamine drop almost entirely below it. From our note The Artificial Hivemind, Revisited (PDF).

We sent the results and the note to Dr Jiang. Her reply, shared on our public submission thread:

"Thank you so much for reaching out and for sharing these results. I'm really glad to see your … system working so well on Infinity-Chat, especially in addressing a challenge that LLMs tend to handle poorly by default. The magnitude of the reduction is particularly striking, and it's also very interesting to see the effect transfer across model providers without modifying the system." Dr Liwei Jiang, author of the Artificial Hivemind study

The Market Is Already Paying for Less

Novelty is not a niche concern; it is becoming a category, and the creative tools already selling it prove the point. The leading one is a model built for creative divergence. It self-reports 7.47 distinctness, with no official benchmark submission and no published utility score at all.

On that basis it has signed major advertising agencies. We read that as market validation: agencies are paying real money for measurably novel output. The system you are about to test holds an officially acknowledged NoveltyBench score above the best published result on both axes, on full-strength frontier models.

Why This Changes the Pitch

Our benchmark result is the first public evidence that the utility tax is not a law of nature: the same system beats the best published results for both distinctness and utility simultaneously.

Behind it sits the verification gauntlet from our full system: a multi-stage adversarial audit in which every claim must survive independent technical, numeric and prior-art attack before it ships. The demo below runs a compressed form of the same discipline. Every answer you will see is scored live by an independent AI judge that played no part in creating it, against the same model asked plainly, with nothing hidden.