Sundae Bar Logo
October 2, 2026

Why we make AI skills compete

By sundae_bar
AI Skills

Every company selling AI tools says its prompts are good. The agents are smart, the instructions carefully engineered, the output reliable. You have read the line so many times it no longer registers.

That is not because it is always false. It is because you have no way to check it. A claim nobody can test carries no information, and AI skill quality has become exactly that kind of claim.

So we built the alternative. Skills compete for a place in our directory, in the open, against a published standard, and the score is printed on the result.

Post Image

The problem with "trust us"

The part that matters is the part you never see

Inside every AI agent is a skill: the written procedure that tells it how to do one job. It decides whether a board pre-read is usable or whether a research brief cites real sources. It is also the part a buyer never gets to inspect. You get a demo, a landing page and the vendor's word.

An unfalsifiable claim stops being heard

"Best-in-class prompts" cannot be proven wrong, so it cannot be proven right either. Buyers have worked this out. They discount the claim to zero and judge on the demo instead, which tells them how the product behaves on one good day.

What we do instead

We replaced the claim with a test anyone can read. It runs in four steps.

1. We publish the job

Each challenge in the sundae_bar Lab names one specific business job, a dataset of realistic cases and a rubric that spells out what a good answer looks like. Before a challenge opens, we test it against skills built on purpose to score well without doing the job. If those can win, the challenge does not open.

2. Independent builders take it on

Builders from outside the company write a skill for that job and submit it. A typical challenge draws dozens of builders and hundreds of attempts at the same task.

3. Every entry is scored the same way

Every submission is scored automatically against the published rubric. Same cases, same rubric, every entry. Nobody at sundae_bar picks the winner, and the results for each entry are published so anyone can check them.

4. Only the winner reaches the directory, with its score

One skill wins. It goes into the directory with its score on the listing. The rest do not.

You can check our homework

This is the part that matters most. We set the job and publish the standard. We do not write the winning skill, and nobody here picks it.

The scoring runs automatically on our systems, so we publish everything it uses: the cases, the rubric and the results for every entry. You do not have to take the score on trust.

That changes what a quality claim is worth. Instead of "our prompts are good", you get something you can check:

  • The job, stated in plain terms
  • The rubric, published with the challenge
  • The score, on the listing
  • The field, how many attempts the winner beat

The numbers so far

  • 15 winning skills published
  • 3,866 scored submissions behind them
  • Every one with its score on its listing

One skill, traced end to end

Launch the Research and Competitive Analyst in Scout and one of the skills doing the work is Citation Audit for Research Briefs.

Hand it a draft brief and the sources behind it. It checks every claim against those sources, grades how well each one is supported, cuts or flags what the sources do not back and logs every change it made. It refuses to treat a reviewer's general knowledge as evidence.

That skill came out of a challenge with 309 competing submissions. Its score, 89, is printed on its listing. You can read exactly what it does before you ever run it.

Why a challenge beats a better copywriter

  • It is measured, not described. A score against a published rubric can be checked. An adjective cannot.
  • It is many attempts, not one team's best guess. Dozens of builders approach the same job differently, across hundreds of attempts, and the strongest approach wins on the evidence.
  • It is built to resist gaming. A skill that only looks good is exactly what each challenge is tested against before it opens.
  • It is hard to copy. A competitor can hire a better writer for their landing page. They cannot retroactively run the challenge.

Why it compounds

Every challenge produces a skill that ships to the directory, where anyone can find and use it for free.

The same skill can then go to work inside the agents you run through Scout. Those agents remember your context, cite their sources and run on a schedule, and challenge winners are built into them.

So each new challenge makes the free directory more useful, and gives the agents another tested skill to draw on.

Where this goes next

Every challenge adds one more job that AI does to a checked standard, and the next challenges are picked from the work people keep asking for.

The next time a vendor tells you their prompts are good, ask to see the score. Ours are printed on the listing of every skill that won. Browse the directory, or start with an agent in Scout.