
AI helped me message 239 people. One in seven said yes.
I tested Jev, a small AI model, on my own LinkedIn outreach: 239 personal messages over two days, and one in seven said yes. Plus GPT-6 Astra vs Opus 5.5.
Round 2 of The JASON test: every new AI model, the same brief.
In September 2026 I tested a small AI model, Jev from TypeSafe, on a real business job: my own LinkedIn outreach. It helped me write to 239 people over two days, at about 90 minutes of sending a day, and one in seven said yes. The same week, OpenAI's GPT-6 Astra and Anthropic's Opus 5.5 both went through my usual website test.
I wanted to test AI on a real business outcome, not a demo, and I picked one use case: growth, meaning new readers from my own network.
How did GPT-6 Astra and Opus 5.5 compare on the same brief?
Same test as always: my CV, one brief, one attempt, no follow-up prompts, each model in its maker's own coding tool. Every number comes from the logs. This was round 2 of the JASON test.
| Opus 5.5 | GPT-6 Astra | |
|---|---|---|
| Run in | Claude Code | Codex |
| Time | 24 min | 36 min 29 s |
| Tokens | 14,935,920 (131,119 output) | 4,962,595 (58,292 output) |
| API calls | 77 | 48 |
| Cost | $6.89 at list price | $13.38 at list price |
| Follow-up prompts | 0 | 0 |
| Lint and build, re-run by me | Pass | Pass |
Opus 5.5 was faster and cheaper. Astra used a third of the tokens but cost about twice as much, because each of its tokens costs more. Both passed first time. Costs are at list price from OpenAI and Anthropic, checked 24 September 2026, with cached tokens at the cached rate. One caveat from the run card: Astra ran in a Codex task that already held earlier portfolio work, so it was not a clean session. Opus 5.5 started fresh.
Both sites are live, so judge the design yourself: the Opus 5.5 build and the GPT-6 Astra build.
What is Jev, and how is it different from a frontier model?
Jev has been all over my feed, and most of it is slop: hype posts, screenshots and big claims, with very little real work behind them. So I gave it the job I avoid most.
Jev is not built to write or build. You ask it a question about a piece of text and it hands back a judgement and a probability. Would this person want to hear from me? Does this line read like a mass email? One small call, made thousands of times, for cents. Frontier models like Astra and Opus are for the one big job; a small model like Jev is for the repeated small ones, and most of the calls in a business look like that. Its first pass over 1,262 LinkedIn profiles took 1,910 calls and cost 10.6 cents.
How did I use a small AI model for LinkedIn outreach?
Given the choice, I will build another feature before I ask anyone for anything. I had 1,262 LinkedIn contacts sitting in an export since June, most of whom I had not spoken to in years. This time I used AI to make that easier.

What it cost me in time:
| Stage | Time |
|---|---|
| Planning | 1 day |
| Building | Half a day |
| Sending | About 90 minutes a day, over two days |
The sending time is measured from the tracker's timestamps: 100 messages across 94 minutes on the first day, 139 across 111 minutes on the second. I sent every message myself.
What did Jev do, and what did I test before sending?
Jev read every profile; I still sorted them by hand. Then it worked all week, and I tested it before I trusted it.

The humbling one: my own draft lost. Jev scored my longer version, with all my reasons in it, as more likely to read like a mass message than a tighter rewrite, for every kind of person I tested. The tighter one went out. And the check for invented familiarity caught one real overclaim: a draft implied we went to the same school. We did not.
How does the outreach tracker work?
It all ran through one page I built: a card per person, Jev's scores on it, one tap to open LinkedIn, and a place to log what came back. Here is a card, with made-up people.
![]()
Sample data: fictional people, the real tracker.
How am I checking whether Jev was right?
Letting Jev pick its favourites and calling the replies a success proves nothing: I would never hear from the people it ranked low. So its predictions were sealed first, and everyone got a message.

It is easy to call something a success when you never said in advance what success was. Most AI pilots I see are decided exactly that way.
What were the results of the outreach?
As of the morning of 24 September 2026. Nobody had yet had a full week to reply, so these are floors, and they are not a verdict on Jev. That waits for the seal to open in October.
| Count | Share | |
|---|---|---|
| Messaged, over two days | 239 | |
| Said yes | 34 | 14.2% |
| Joined the newsletter, of the 234 I asked | 32 | 13.7% |
| Said no | 2 | 0.8% |
| No reply yet | 203 |
People I know joined at 21.8% (22 of 101). People I had never spoken to joined at 7.5% (10 of 133). The typical yes arrived within a few hours.
What surprised me most was my own gut. Before each message I guessed how likely that person was to reply, and I said Low for 233 of 239. One in seven of them said yes. And of the 116 people I skipped, I chose not to contact 99: more than a quarter of everyone I got to, on a list I had already sorted by hand.
How does that compare with LinkedIn outreach benchmarks?
I watched one number: of the people I asked, how many did the thing I asked for?

13.7% gave me their email, from one message and no follow-ups. Published benchmarks for any reply at all sit between 10% and 12%: 12.2% for messages to existing connections in the Belkins LinkedIn outreach study (2025), and 10.4% across 6,730,447 messages in Expandi's 2026 benchmarks, follow-ups included. It is not a controlled comparison, and people I know did most of the lifting.
Among people I had never spoken to, the group Jev rated most likely to join did join more often: 9.7% (3 of 31), against 4.2% (1 of 24) for a group chosen at random. Right direction, but it is four people, and I skipped more of the random group, so it proves nothing yet. October will.
What should a business take from this?
My outcome was growth. AI made it affordable to treat 239 people as 239 people, each with the right message and the right example of my work, in about 90 minutes a day.
The list is an asset now, too. Going through it meant labelling my own network: who I know, who I have never spoken to, who to leave alone, who to come back to. Those labels stay private, in my AI OS, the same set of folders my voice notes land in. The next time I ask AI about my network, it starts with context instead of a blank page.
The new models will not grow anything on their own. They make the personal version affordable at scale, and they leave you with data you did not have before. Pick one number that means someone acted, write it down before you start, and check it against a benchmark. Then keep what you learn where your AI can read it.
Common questions
What is Jev? A small AI model from TypeSafe that does not write text. You ask it a question about a piece of text and it returns a judgement and a probability, cheap enough to ask thousands of times.
How much did it cost to score 1,262 LinkedIn profiles with AI? 10.6 cents, for Jev's first pass of 1,910 calls. The whole week took 14,221 calls.
What is a good reply rate for a LinkedIn message? Published benchmarks for any reply sit between 10% and 12%: 12.2% for messages to existing connections (Belkins, 2025) and 10.4% across 6.7 million messages (Expandi, 2025 to 2026, follow-ups included).
Is GPT-6 Astra or Opus 5.5 better for building a website? On one brief, Opus 5.5 was faster at 24 minutes and cheaper at $6.89. GPT-6 Astra used a third of the tokens but cost $13.38. Neither was declared the winner on design, and both sites are live to compare.
Every new model gets the same brief in the JASON test, and the model routing behind all of this is in what kind of model is GPT-6 Astra. If you would rather build a system than read about one, the free course comes with the newsletter.
This started life as Build Notes Nº 06, in September 2026.
Get the next one
Subscribe and get the free course, Build Your First AI Operating System. Then a few emails a month on what I build, including what didn't work. See the course