Learn how HiPeople is impacting the world’s leading hiring teams in a live demo session. Register

TypeSafe AI and what a new kind of model could mean for hiring

TypeSafe AI and what a new kind of model could mean for hiring

Since last Thursday, our engineering team has been evaluating TypeSafe AI in secure environments.

We’re seeing significant potential to make parts of our agents running hiring end-to-end more precise, faster, and more affordable.

What makes TypeSafe extremely interesting is that it is a fundamental different model.

Meet TypeSafe

TypeSafe AI is a new AI company building models for software decision-making rather than conversation. TypeSafe was founded by Diogo Almeida, who worked on RLHF and InstructGPT at OpenAI, alongside former Meta/FAIR research engineer Sasha Sheng and repeat AI founder Erik Gafni, and is backed by a $40M seed round led by DCVC.

Its first public model, Jev, launched September 15, 2026.

For those not-technical: the name TypeSafe comes from a concept in software engineering. A “type” defines what kind of value a system expects: a number, a date, a category, a yes or no. “Type-safe” software makes sure the value returned is actually one the system knows how to work with.

TypeSafe brings this idea to AI.

Instead of asking a model to generate an open-ended response and then turning that response into something software can act on, the possible outcomes are defined upfront.

The model makes a decision within that structure and returns probabilities for the possible outcomes.

That sounds like a small distinction.

It isn’t.

Generating answers and making decisions

Large language models are designed to generate language. Give an LLM a question and it predicts the next token, then the next, until it has produced an answer. This makes LLMs remarkably general. They can analyze a résumé, summarize an interview, write an email, reason through conflicting information, or explain why a candidate could be a strong fit.

But much of the work performed by an AI agent looks different.

An agent constantly needs to make decisions.

Does this candidate have the required experience?

Is this skill supported by the available evidence?

Does this application satisfy a requirement?

Which category does this information belong to?

How confident are we?

For these tasks, generating language is often only a means to an end. What the software ultimately needs is a decision.

Jev is built around that idea.

A traditional LLM

Context → Generate response → Interpret response → Validate → Determine confidence → Decide

Jev

Context → Probabilistic decision

Instead of generating an answer that a system then has to interpret, Jev returns a structured decision with probabilities built in.

Why this matters

Getting an LLM to produce one answer is easy.

Getting it to produce decisions that are consistently reliable enough for software to act on is a much harder engineering problem.

Today, sophisticated AI systems use a range of techniques to get there.

A task can be broken into multiple smaller model calls. The same question can be sampled several times and the results compared. Multiple models can evaluate the same evidence. Another model can act as a judge. Outputs can be verified before the system proceeds. These techniques can work extremely well. They also require more calls, more tokens, more time, and more infrastructure.

Jev approaches the problem differently. Probabilities aren’t something added around the model. They are part of its output. One call can return both the decision and how probability is distributed across the possible outcomes. This can make a class of AI tasks significantly simpler, and potentially much faster and more affordable.

Faster and cheaper doesn’t just mean faster and cheaper

This is the part we find particularly interesting. The obvious benefit of reducing the compute required for a decision is cost and speed. The more important benefit is what you can do with the compute you get back.

You can evaluate more evidence, make assessments more granular, add safeguards, and run additional checks. And you can reserve the most powerful reasoning models for the parts of a workflow where deep reasoning actually makes a difference.

In other words, making individual decisions cheaper creates room to make the overall system more intelligent.

What this could mean for hiring

Hiring is full of probabilistic decisions. A candidate rarely simply “has” or “doesn’t have” a skill. There is evidence, context, strength, ambiguity, and confidence. The same is true when evaluating experience, qualifications, role requirements, seniority, domain expertise, and many of the other signals involved in understanding a candidate.

These are exactly the kinds of decisions an agent running hiring end-to-end needs to make at scale. Today, we already use different models for different parts of that journey: some tasks benefit from deep reasoning, others from fast extraction, generation, or communication. Models like Jev could add another specialized capability: fast, structured, probabilistic evaluation.

For the right tasks, that could mean more precise assessments in less time and at lower cost. And that efficiency compounds: if evaluating one signal becomes cheaper, an agent can evaluate more signals; if an assessment becomes faster, it can perform additional verification without making the recruiter wait. The result is room to make the overall system more intelligent.

What this doesn’t mean

It doesn’t mean LLMs are going away. Jev isn’t designed to replace models that excel at understanding complex context, reasoning through unfamiliar problems, or generating language. Those capabilities remain essential.

Instead, we think AI systems will increasingly combine different kinds of models for different kinds of work. A reasoning model can make sense of a complex situation, a probabilistic model can make thousands of tightly defined evaluations, and a language model can turn the results into an explanation someone actually wants to read. The agent brings them together.

The important question isn’t which model is best. It’s which intelligence is best for the task at hand.

What’s next at HiPeople

For now, we keep testing. We’re evaluating Jev across hiring tasks where its architecture could create a meaningful advantage, looking at precision, calibration, latency, cost, and how it performs across the edge cases that appear in real hiring workflows. If we’re confident it improves the product, we’ll introduce it where it makes sense.

We’ll do so within the data residency requirements our customers have configured in HiPeople. TypeSafe is currently available with US processing, so customers requiring sole EU residency will take a little longer to benefit. If TypeSafe becomes a production subprocessor for applicable customers, we’ll update our subprocessor information accordingly.

More broadly, TypeSafe is another example of how quickly the underlying intelligence available to us is evolving. Our job is to turn that intelligence into something useful for hiring: choosing the right models for the right tasks, giving them the right context, building the right guardrails and compliance infrastructure around them, continuously improving them with feedback, and putting all of it behind a product people actually want to use.

The models will keep changing. Our goal stays the same: apply the best available intelligence to hiring, so people can spend less time operating hiring processes and more time on the human decisions that matter.

Discover how HiPeople makes hiring more human

HiPeople helps the best hiring teams find incredible talent that sets their organizations apart.