Skip to content
Put your prompt on trial

Find the prompt that gets the best answer, and prove it.

PromptJury treats your prompt like an experiment. Define what a great answer looks like, test every prompt and model on realistic cases, and keep the one that scores highest.

No card. You bring an OpenRouter key, and you see the estimated model cost before anything runs.
Test case 14 of 40 · hard

“The vase arrived cracked. It's been 40 days because I was travelling. Can I get my money back?”

✓ Best answer8.6 / 10

I'm sorry it arrived broken. Damaged items are always refunded, whatever the date, so you'll get the full amount back. Reply with a photo of the vase and I'll process it today.

prompt v3 · openai/gpt-4.1
First draft5.2 / 10

Unfortunately our refund window is 30 days, so we can offer store credit instead.

prompt v1 · same model

Why v1 lost: “Misses that damaged items are always refunded.” 3 of 3 graders, graded blind.

Sample result, for illustration
Most prompts are tuned by feel.

You change a line, read three outputs, and decide it's better. That's a hunch, not a result. It holds until a customer asks something you didn't try. PromptJury writes the test first, runs the tries you didn't have time for, and keeps the receipts.

The scientific method, for prompts

Six steps, in one sitting. Each starts with a draft you can accept or change, so you never face a blank form, and none of it needs code.

  1. State the goal.

    Two sentences on what your AI should do and who reads the answer.

  2. Write competing prompts.

    Several models each draft one, then critique each other's. You start with rivals to test, not a single guess.

  3. Define success before you test.

    Approve what a great answer looks like, dimension by dimension, with a scale for each. The test exists before the prompt is judged.

  4. Gather test cases.

    Realistic inputs are written for you, including the hard and awkward ones. Or upload your own.

  5. Run a fair experiment.

    Every prompt and every model answers every test case. A separate panel grades each answer blind.

  6. Keep the winner, then improve it.

    See the best prompt, the best answer and its score. Each model gets a suggested fix; re-test it and watch the score move.

Why you can trust the score

“One model grading its own work is a student marking their own exam.” So the grading is a controlled experiment.

  • Graded blindGraders never see which model or prompt wrote an answer.
  • No self-gradingA model never scores its own answers.
  • Several graders, median scoreOne generous or harsh grader can't swing the result.
  • A range on every scoreEach score comes with its likely range, from resampling the test cases.
  • Honest about tiesWhen the top two are within the margin of error, we say “too close to call”.
  • Every number has a receiptOne click from any score to the answer and each grader's reason.
One answer, four jurors, names hidden
  • Juror A8 out of 10
  • Juror B9 out of 10
  • Juror C7 out of 10
  • Juror D8 out of 10
Median score8.0
Sample data

Improve, re-test, repeat

The verdict says what each prompt failed to tell each model. Apply a fix, run the same test cases again, and keep the change only if the score goes up. Every version has to earn its place.

Saved cases become regression tests. Schedule a retrial and PromptJury re-runs them when models change, so you hear about a drop before your customers do.

56789106.47.68.6
  1. v1 · First draftAnswers politely, skips the next step
  2. v2 · DeliberationAdds the 30-day rule; still vague past it
  3. v3 · Suggested fixOffers store credit past 30 days; ends with one action
Same 40 test cases each time. Lines show each score's likely range.
Sample data

A sample result

An illustrative experiment: support replies for a home goods shop, five models, 40 test cases, three graders. Every model's score sits next to its cost, so you can see the best answer and the cheapest one that is nearly as good. Pick a model to see its numbers.

Browse published results →
Score out of 10 ↑ · cost per 1,000 calls →
gpt-4.1claude-sonnet-4.5claude-haiku-4.5gemini-2.5-flashmistral-small
✓ Winneropenai/gpt-4.1
8.6
/ 10, range 8.2 to 8.9
$1.55
per 1,000 calls
A juror's reason, on the day-45 refund case

“Holds the 30-day policy and offers store credit instead.”

Why your own key

You pay model makers directly

You connect your OpenRouter account, so calls go through it at the model makers' own prices.

No markup

We add nothing to model costs. Your plan pays for PromptJury; your key pays for the calls.

An estimate before every run

Nothing that spends money starts until you've seen the estimated cost and confirmed it.

Who it is for

Solo builders

“Tell me which prompt and model to ship, and show me why.”

You can wire an API call. You don't have time to build a test harness. Get a tested prompt in one sitting.

Small agencies

“I have nothing to back up my model choice.”

Hand clients a report with your logo on it: the method, the test cases and the scores behind your choice.

Business owners

“I only hear about it when it's bad.”

Schedule a retrial and get a plain pass or fail by email, before a customer notices.

  • Free$0 1 case, 3 trials a month
  • Trial Pass$19 once10 trials over 14 days, PDF reports
  • SoloMost builders$29 a month30 trials a month, monthly retrials
  • Pro$99 a monthWhite-label reports, 5 seats
You pay your own model costs through OpenRouter. We add nothing.

Fair questions

What makes this scientific rather than another playground?

You fix the definition of a good answer before anything is scored, every prompt and model faces the same test cases, grading is blind, and each score comes with a range. Change one thing, re-run the same test, and you can see whether it actually helped.

Will it find the perfect prompt?

It finds the best prompt among the ones tested, against the standard you approved, and shows how sure that result is. Then it suggests what to try next. Nothing can promise perfect; this gets you measurably better, with proof.

Isn't AI grading AI circular?

You approve the grading guide, and the verdict shows sample answers so you can check the scores against your own reading.

Won't a model favor answers from its own maker?

It can, which is why the jury mixes makers and never sees which model wrote an answer. No model grades itself. You can see every juror's score and reason, and we flag dimensions where they disagreed.

Can't I do this in a spreadsheet?

You can. It takes a day per prompt, and you will not do it again when the next model ships.

Will this run up my API bill?

Every trial shows an estimate first, and nothing that spends money runs until you confirm it. You can also cap spending on the key you connect.

My outputs are subjective. Does this still work?

That is the case PromptJury is built for. Subjective quality is graded on dimensions you approve.

Do I need an OpenRouter account?

Yes. It takes a couple of minutes, and one key reaches models from every major maker. We walk you through connecting it on your first case.

Can I use my own test cases?

Yes. Upload a CSV and map its columns, or edit the cases we generate. Most people do a bit of both.

Do you train on my prompts or data?

No. Your prompts, test cases and answers are used only to run your trials, and you can delete a case at any time.

What if the top two are too close to call?

We say so in plain words instead of crowning a winner. The verdict shows the range for each model, and the cheaper one is usually the sensible pick.

Stop guessing. Test your way to the best prompt.

Your first case is free, with no card. If the result does not tell you something you did not know, you have lost half an hour.