You can set up your own critics in a general AI chatbot — a custom persona, a saved project, a stack of system prompts. For a casual gut-check, do it. But the persona is the easy tenth. The nine-tenths that make a critique trustworthy instead of merely plausible is the machinery underneath: findings forced to quote your own words, a cross-examination pass that kills the weak ones, a score computed by formula, and a coded failure taxonomy calibrated against plans that actually died. That is what BadMars is — and it is not a prompt.
1 free Mission a month · no card · verdict in ten
| Building it yourself | BadMars | |
|---|---|---|
| What you get | A persona that role-plays a critic | An engineered adversarial procedure |
| Findings tied to your words | Only when the model feels like it | Mandatory — every finding quotes your text |
| Weak findings removed | No — whatever it says, you keep | Yes — a separate cross-examination pass |
| Score | A number by mood, different each run | A fixed formula, the same standard every time |
| Grounding | The model's general priors | A 53-code failure taxonomy calibrated on real outcomes |
| Politeness drift | Slides back to agreeable across a session | Fought structurally, by design |
| Memory | Starts from zero each time | Versioned Programs · Score trajectory · re-runs |
| To build it well | Weeks of prompt work across many fields, then upkeep | €39 and ten minutes |
You want a quick, casual second opinion and you know its limits — that it may invent problems, miss the ones that matter, and cheer you on regardless. For a five-minute gut-check, that is enough.
The decision actually matters and the critique has to be trustworthy, not just confident. Anchoring, cross-examination, a formula score, and a calibrated taxonomy are the difference between "this sounds smart" and "this is where your plan breaks."
You can build the personas — that part is easy. What you cannot easily build is the procedure that makes their output trustworthy: forcing every finding to quote your own words, running a second pass that removes the weak ones, scoring by formula so the number does not move with the model's mood, and grounding it all in a failure taxonomy tuned against plans that really failed. Without that, a custom persona produces plausible-sounding critique with nothing holding it to the truth.
The prompts are the smallest part. The engine is versioned specialist briefings, a coded 53-code failure taxonomy applied from the first report, a cross-examination stage, a formula score, and calibration from real outcomes over time. A weekend clones a prompt; it does not clone that.
Not the model — you can rent that anywhere. You are paying for the adversarial machinery around it: thirteen specialist briefings built and maintained across finance, competition, operations, regulation and more, the procedure that keeps them honest, and the calibration that keeps improving — for €39 instead of the weeks it would take to approximate, badly.