Jev for GTM
Most GTM teams use a large language model for jobs that don’t need one. Scoring an account, tagging an industry, deciding if a reply is interested - those are decisions, not writing. Paying a model that writes essays to answer “yes or no” gets slow and expensive at scale.
Jev is built for exactly those decisions.
Related: For the research-then-write setup we use for personalized emails, see Apollo.io AI research for personalized outreach. Jev handles the step before that - deciding who is worth researching at all.
What Is Jev?
Jev is the first “System One” model from TypeSafe. It doesn’t write text. You give it the state of something (a company profile, a website, a reply, a CRM record) and a set of typed questions. It returns answers your code can use directly:
- Yes/no questions come back as a probability.
- Multiple choice comes back as the pick plus a probability for every option.
- Scores come back as a position on a scale you define.
Every answer carries a confidence score, so your code knows when to trust it and when to hand the case to a person or a bigger model.
The pricing is what changes the math: $0.042 per million input tokens, and output is free. Answers come back in well under a second.
Why Jev Fits GTM Work
GTM data is messy. A company’s website, its Apollo.io record, its tech stack, its job posts, and its LinkedIn page all say slightly different things. Deciding what that pile of signals adds up to is a judgment call, and it has to be made thousands of times.
That’s where Jev earns its place. It prioritizes, qualifies, and scores at a cost low enough to run on every account you have, and fast enough to run again whenever the data changes.
Any good LLM can make the same calls with similar accuracy. The difference is cost and speed. A frontier model reading every account in a large list burns millions of tokens and takes hours. Jev does the same pass for pocket change in minutes.
GTM Use Cases for Jev
Account scoring
- Score every account in your CRM against your ICP on several dimensions: fit, timing, size, tech stack.
- Keep the weights in your own code, so changing what “tier A” means never requires re-running the model.
Account classification
- Sort accounts into verticals, business models, or buyer types from messy website and firmographic text.
- Catch what keyword filters miss, like a company that never says “payments” but obviously takes them.
Lead qualification
- Check inbound leads against your disqualifiers (agency, student, competitor, wrong geography) before a rep sees them.
- Route the uncertain ones to a human instead of guessing.
Reply triage
- Label replies as interested, not now, wrong person, or unsubscribe.
- Get warm replies to the rep fast and keep the rest out of their inbox.
Signal and intent filtering
- Decide whether a job post, funding round, or tech change is actually relevant to your offer.
- Turn hundreds of raw signals into a short list worth acting on.
Checking AI output before it ships
- Verify that a personalized line is supported by the page it came from.
- Block anything that states a fact the source doesn’t back up.
The pattern is the same every time: Jev makes the call, code applies the rules, and a writing model like Claude only gets involved where words are needed.
Jev Account Scoring: Re-Scoring 12,700 Accounts
Here’s how we used it on a real client list.
The client sells embedded payments to software companies. Their Apollo.io instance had 12,733 accounts marked “Not a Fit” by an AI prompt that ran a year earlier, every one of them parked in Do Not Prospect.
The list included 92 point of sale companies, 131 medical practice management platforms, and 94 property management tools, all tagged “no payments need.” One old AI reason described a cloud infrastructure tool as software whose users “accept digital payments.” It made that up.
Why Did the Old Qualification Prompt Get It Wrong?
The old setup asked one big question per account and told the model to say no when unsure. That breaks in predictable ways:
- It writes a reason even with no evidence. That’s how the made-up answers got in.
- “When unclear, say no” throws away good accounts. A vet practice management system takes payments at checkout even if its homepage never says so.
- One big question hides what went wrong. You can’t tell if it misread the company, the product, or the rule.
How We Built the Jev Account Scoring Pipeline
Step 1: Gather evidence for free
- Pulled every account from the Apollo.io list (account searches don’t cost credits).
- Fetched each company’s homepage plus up to three pages about payments, pricing, or features. About 12,000 sites took 15 minutes.
- Built a short evidence card per company: title, description, headings, menu, sentences that mention payments, and any payment processor named on the site.
Step 2: Ask Jev five narrow questions per company
- Does this company sell software, or is it an agency, a hardware maker, or the business itself?
- Do its customers collect money from their own customers through it?
- Ignoring the website, is this the kind of software that normally takes payments?
- Does a disqualifier apply (internal tool, services firm, payment processor, excluded industry)?
- Which vertical does it serve?
All five go in one request, so each company is one call.
Step 3: Test Jev against Claude
Before trusting it, we had Claude judge 300 of the same accounts blind and compared:
- Jev’s confident “not a fit” calls matched Claude 168 out of 168 times. It never threw out a company Claude considered a fit.
- With one cutoff applied to every answer, the two models agreed on 94% of accounts.
- Most disagreements were companies Jev leaned toward fit and Claude rejected, which is the safe direction because Claude reviews those anyway.
Step 4: Score the survivors with Jev again
Jev ruled out 9,010 of the 11,748 companies it checked. Instead of handing the remaining 2,738 to Claude, we ran a second, deeper Jev pass with more evidence per company and graded questions:
- How strong a prospect is this, on a five-level scale?
- Is taking payments core to the product, or a side feature?
- Does the site say so explicitly?
- Are the users small and mid-size businesses?
- Is this company itself a payments company?
Code turned those answers into a 0 to 100 fit score. Jev also picked the single best payment sentence from each site, which gave us a word-for-word quote without a writing model. That pass took 30 seconds and cost 20 cents.
Against Claude’s blind judgments, 91% of companies scoring 80 or higher were real fits. That became the Fit list: 1,072 companies.
Step 5: Use Claude only on the best accounts
Claude read only those 1,072 (minus 28 already on do-not-contact or partner lists) and did what Jev can’t: confirmed each one, wrote a specific one-line reason, and saved a research note with who pays whom, payment features, the current processor, and one observation quoted from the company’s own site. A script then checked every quote against the saved web pages.
How Much Cheaper Is Jev Account Scoring Than an LLM?
We measured both on the same data.
| Claude for every account | Jev for scoring, Claude for the best accounts | |
|---|---|---|
| Companies Claude reads | 11,748 | 1,344 (300-account test + 1,044 fits) |
| Claude tokens | ~16.7 million to judge, more for research notes | ~2.9 million, research notes included |
| Jev tokens | none | 25.5 million |
| Jev cost | $0 | $1.07 for both passes |
| Time to score every account | Hours of batched agents | About 2.5 minutes |
Jev actually used more tokens per company than Claude: about 1,800 against about 1,400, because the questions ride along with every request. It doesn’t matter. At $0.042 per million with free output, scoring the entire list twice cost about a dollar, and Claude’s workload dropped by more than 80%, with the research notes included.
What Did Jev Account Scoring Find?
Out of 12,733 accounts the old prompt had written off:
- 1,015 are confirmed fits. 1,007 of them had been marked “no payments need” by the old prompt.
- 28 more fits were held back because they were already on blacklist, partner, or active-deal lists.
- Claude disagreed with Jev on only 29 of the 1,044 top-scored companies, about 3%. Almost all were resellers, a few payment processors, and two business banks.
- 1,040 of 1,044 quotes were found word for word on the company’s own site. The four that weren’t got dropped.
The fits are exactly what the client sells into: retail and POS (137), property management and rentals (101), restaurants and hospitality (93), medical and dental (71), events and sports (69), fitness and membership (59), plus education, field service, auto, nonprofits, and veterinary.
The payment picture is useful too. 159 already run their own branded payments, 329 name a third-party processor (Stripe alone in 140+), and 198 show no payments at all. Each of those is a different opening line, and the research notes already say which applies.
How We’re Using Jev Across GTM Works
Account scoring was the first job. These are the places we’re rolling Jev into next, each one a decision we currently pay a bigger model (or a brittle keyword rule) to make:
- Reply triage: every reply to a client’s Apollo.io sequences gets sorted into interested, not now, wrong person, or unsubscribe so warm leads reach the rep fast. Jev makes that call for a fraction of a cent and sends only the unclear replies to Claude.
- Positive reply reporting: our monthly reports sort every email and LinkedIn reply into positive, referral, or negative before we write recommendations. Jev does the sorting so Claude only writes the insights.
- Personalized openers: before an AI-written opener goes into a sequence, every claim gets checked against the page it came from. Jev runs the first check and flags anything the source doesn’t support.
- LinkedIn managed outreach: when we refill campaigns from Apollo.io, Jev confirms each person actually works at the company the campaign was researched for, which stops the wrong-company messages that LinkedIn profiles with several jobs cause.
- Signal-based outreach and custom lead engines: Jev reads job posts, website changes, and tech signals and decides which ones actually matter for the offer, so reps act on a short list instead of a raw feed.
- Exhibitor and sponsor lead generation: Jev classifies event organizers and exhibitors from their own websites instead of keyword rules, which miss anything phrased differently.
- Domain matching: when we match a company list to websites, Jev confirms each homepage really belongs to that company. In testing it scored the right sites at 0.97 and look-alikes (Jersey Mike’s for “Jersey Watch”) at 0.09.
Same pattern every time: Jev makes the call, code applies the rules, and Claude only writes where words are needed.
Do You Want Jev for GTM Set Up for Your Company?
We set up Jev account scoring, reply triage, and the rest of this stack on top of your Apollo.io account, as part of our Apollo.io setup and managed services. If your CRM has thousands of accounts nobody has looked at in a year, we’ll show you which ones are worth working.
Contact us to get Jev for GTM running for your team.