Making Decisions with Jev and Goose
GovernmentArchitectureSecurityTypeScriptOpen SourceAI

Making Decisions with Jev and Goose

·Neil ChaudhuriNeil Chaudhuri on LinkedIn

Most of the work your business runs on comes down to decisions. Your teams route support tickets, approve or flag invoices, and triage security alerts. Your AI agents spend their days on the same kind of work, and nearly every one of those tasks has a short list of acceptable answers. None of them needs a model that can write poems from the perspective of Voldemort in Harry Potter.

Many teams send those decisions to the same frontier LLM that drafts their emails. An LLM writes its answer one token at a time, so even a one-word label arrives as generated text that your code must parse and validate before it can act. You wait for and pay for every one of those tokens.

A decision model answers the question directly. Jev, the first decision model from TypeSafe, takes your application’s state along with a set of typed questions and returns typed answers with probabilities in a fraction of a second. You can run Jev from Goose, the open source agent that Block contributed to the Agentic AI Foundation (AAIF), with an LLM that orchestrates the work flow and interacts with you through natural language. I believe that pairing beats either model on its own. The LLM speaks natural language with the people who use the agent, and the decision model makes cheap, fast, calibrated calls for the software behind it.

The Power of a Decision Model

A recent Hugging Face guide to choosing between Jev and an LLM makes the distinction concrete. An LLM generates content for people to read, and even its structured output remains generated text that your code must check for missing fields and formatting drift. A decision model starts from an answer space you define and returns a yes-or-no probability, choice, or a score for software to act on. The guide offers a simple test. If you can list the possible answers in advance and your code will route, rank, or block on the result, use a decision model.

TypeSafe calls the building blocks of that answer space primitives, and each one pairs a question type with a typed answer:

  • Noul answers a yes-or-no question with the probability of yes. “Is the customer satisfied?”
  • Choice picks one option from a set you define and returns the top option, the probability of every option, and a confidence value. “Is the customer most likely to buy drinks, appetizers, meals, or desserts?”
  • Score places the state on an ordered scale of levels you describe and returns a probability-weighted position on that scale. “On a scale from Kid Rock to Taylor Swift, how much did the customer like this song?”

One request can ask several questions about the same state, and Jev answers each one independently under an identifier you choose. Your code can ask for a department, an urgency level, and a refund flag in a single call and then tune the threshold for each answer without touching the others.

TypeSafe calls Jev a System One model, a name that borrows from the fast, automatic System 1 in Daniel Kahneman’s account of human judgment. TypeSafe trains these models with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD) to provide the kind of judgment an expert makes in about a second. The home page argues that the Reinforcement Learning from Human Feedback (RLHF) behind chat models makes them overconfident. Calibration means the probabilities track real outcomes across many predictions. TypeSafe’s documentation also cautions that calibration cannot guarantee any single answer, so your code should weigh each probability as evidence before it acts.

The Need for Speed

Over the last few years, we have delegated decisions to frontier models with type safety optional although I am personally relentless about type safety. It was the best we could do even if we know deep down the judgments LLMs give us and the probabilities they may ascribe to their judgments are more or less confident fiction.

The Hugging Face guide suggests a decision model whenever a workflow needs many small judgments, routing by probability, or a decision interface that holds steady at high volume. I completely agree because decision models are built for that while LLMs are not.

TypeSafe’s workflow evaluations put numbers on that choice. The evaluations cover 705 cases across four automation workflows, from invoice fraud triage to security alert triage, and every model runs at its provider’s default reasoning setting. Across all four workflows, Jev matched the 67.8 percent average accuracy of Claude Sonnet 5 for $0.0004 per case instead of $0.1174. Jev also finished each case in 0.4 seconds against 78.1 seconds for Claude Sonnet 5. At those prices, your agent can run more than 250 cases through Jev for the cost of one case on Claude Sonnet 5.

The same evaluations make a second point that matters even more. Every model that TypeSafe tested in both modes scored higher as a structured workflow of small decisions than as one standalone prompt. Claude Haiku 4.5 alone jumped from 18.1 percent to 53.6 percent. Your agents gain accuracy as soon as you break their work into decisions, and a decision model makes each of those decisions cheap.

TypeSafe publishes the harness on GitHub and the datasets on Hugging Face, so your team can rerun the comparison on your own cases before committing. My own measurements are consistent with theirs. The two requests behind my sources sought example below made the round trip from my Mac to TypeSafe in 169 and 184 milliseconds, network included. At that speed, a check before each tool call adds less than a fifth of a second to your agent’s work, which I consider a bargain for a calibrated answer.

Teaching Goose to Use Jev

I cannot stress enough how much AI demands interoperability so we can enjoy freedom of choice with our tools, and Goose shines here. It also supports Agent Skills, the open standard for packaging instructions, domain expertise, and resources that an agent loads on demand, and it discovers global skills in ~/.agents/skills. A skill you write for one agent carries over to Goose. This is the beauty of standards, which I am also relentless about.

TypeSafe publishes an agent skill that teaches a coding agent the three question types, their architectural patterns, and their practices for evaluating results. Any agent that supports the standard can install it with one command.

npx skills add typesafe-ai/skills --skill typesafe-ai

I made two changes before I let Goose use it.

Bundled References

TypeSafe’s skill points the agent to its llms.txt documentation index and has it fetch pages such as the API reference and the guide for each primitive from their site. I replaced those links with a references folder inside the skill, a layout for progressive disclosure that the Agent Skills specification also recommends for documentation. The folder holds a copy of TypeSafe’s API reference and a condensed guide to writing questions. The skill tells Goose to read those files and to fetch nothing without asking me first. Goose now works from documentation I reviewed instead of whatever a web page somewhere says at the moment it loads.

That difference matters for security. Every page an agent fetches can carry a prompt injection, the first risk in the OWASP Top 10 for LLM Applications. I have warned before that borrowed agent content is a vulnerability. The bundled folder also saves a network round trip on every task, and in some of my environments it saves an approval prompt too.

Bundled references do offer one risk. TypeSafe’s own troubleshooting guide warns that a stale skill can lead an agent to invent request or response fields, and my copy starts to age the day TypeSafe changes its API. Every Jev response names the model that produced it, jev-1.13.0 as I write this, so a change in that field tells you it is time to compare the folder against TypeSafe’s documentation. This is your regular reminder that we still have to wrestle with breaking changes in the age of AI.

A Key from the Keychain

TypeSafe’s quickstart expects the API key in the TYPESAFE_API_KEY environment variable. An environment variable travels to every process the shell starts, and an agent that prints its environment can copy the key straight into a transcript. My version of the skill tells Goose to read the key from the macOS keychain through the Computer Controller extension, which ships with Goose and automates macOS applications and system tasks. The keychain encrypts the key at rest, and the skill has Goose read the key and send the request in one command so the value never lands in the transcript.

Storing the key is easy. With -w at the end, security prompts for the value and keeps it out of your shell history.

security add-generic-password -a "$USER" -s "TYPESAFE_API_KEY" -w

The call reads the key and uses it in the same step. Piping the header into curl -K - also keeps the key out of the process list.

KEY="$(security find-generic-password -a "$USER" -s "TYPESAFE_API_KEY" -w)"
printf 'header = "Authorization: Bearer %s"\n' "$KEY" |
  curl -sS -K - https://api.typesafe.ai/v1/systemone \
    -H "Content-Type: application/json" \
    --data-binary @request.json
unset KEY

The same pattern works with any secret store your platform offers, and it keeps every credential your agents use out of their transcripts.

A Trivial First Decision

A trivial example shows how the two models split the work. I asked Goose, running Gemini 3.1 Pro, to judge the sentiment of one sentence about my favorite baseball team.

Using the typesafe skill tell me if this is a positive sentiment: "My favorite
baseball team is in the top 10 in team OPS in Major League Baseball."

Goose loaded the typesafe-ai skill and its references/api.md file, ran one curl command against the TypeSafe API, and returned this response:

Using the `typesafe-ai` skill, I evaluated the statement with the `jev-latest` model.

**Yes, this is a positive sentiment.**

The model returned a `noul` (probability) of **0.96**, indicating a 96% certainty that the statement expresses a positive, happy, or proud sentiment. Being in the top 10 for team OPS (On-base Plus Slugging) is a strong offensive statistic in baseball, making it a positive statement to make about a favorite team!

The whole exchange finished within the same minute, and Goose reported $0.07 of Gemini usage for it. Gemini also explained that a Top 10 rank in OPS marks a strong offense, which is true. That is context that Jev may not know and would never tell me anyway because it returns only typed values.

This might be a silly example, but it gives you a sense of how decision models and LLMs complement each other so well.

Gemini Talks, Jev Decides

That exchange gave each model the job it does best. Gemini handled every step that needed language. It understood a request in natural language. It then wrote a valid API request from the bundled reference and explained Jev’s answer back to me in natural language. Jev handled the one step that needed a judgment. The result is a cheap, fast decision that an LLM orchestrates and explains in a manner I can easily understand.

The typed answer matters as much as the explanation because it lets the same judgment move from a conversation into code. Gemini translates natural language into the types Jev needs and then interprets the types in Jev’s response to the language I need.

Here is a more interesting example.

Using the typesafe skill, tell me what the sender of this contact form message wants from
Vidya and whether the message names a deadline: "Hi Neil, our team runs a 15-year-old Java
application on servers we have to vacate next year, and nobody left on the team knows how it
works. Could Vidya help us move it to the cloud before our data center lease ends in June?"

Gemini designed the questions itself and turned that prompt into this request.

{
  "model": "jev-latest",
  "state": "Hi Neil, our team runs a 15-year-old Java application on servers we have to vacate next year, and nobody left on the team knows how it works. Could Vidya help us move it to the cloud before our data center lease ends in June?",
  "questions": {
    "service_requested": {
      "type": "choice",
      "instructions": "What primary service is the sender requesting from Vidya?",
      "criteria": {
        "cloud_migration": "Moving an existing legacy application or system to cloud infrastructure.",
        "new_development": "Building a brand new software application from scratch.",
        "team_training": "Training or upskilling the existing team on new technologies.",
        "unclear": "The requested service is not clear or is something else."
      }
    },
    "has_deadline": {
      "type": "noul",
      "instructions": "Does the message explicitly name a deadline or time constraint?",
      "criteria": {
        "true": "The sender states a specific month, timeline, or event by which the work must be completed.",
        "false": "The sender does not state any time constraints or deadlines."
      }
    }
  }
}

Jev returned typed answers under the identifiers that Gemini chose.

{
  "model": "jev-1.13.0",
  "answers": {
    "service_requested": {
      "type": "choice",
      "choice": "cloud_migration",
      "confidence": 1.0,
      "probabilities": {
        "unclear": 0.0,
        "new_development": 0.0,
        "cloud_migration": 1.0,
        "team_training": 0.0
      }
    },
    "has_deadline": {
      "type": "noul",
      "noul": 0.98
    }
  },
  "usage": {
    "input_tokens": 502,
    "output_tokens": 78
  }
}

Gemini then translated those answers back into natural language.

Using the `typesafe-ai` skill, here is the structured evaluation of the contact form message:

- **What they want:** The model determined with **100% confidence** (`choice: "cloud_migration"`) that the sender wants help moving their legacy application to cloud infrastructure.
- **Is there a deadline: Yes**. The model gave a **0.98** probability (`noul: 0.98`) that the message explicitly names a deadline or time constraint. (Specifically, the text mentions having to vacate servers next year and the lease ending in June).

Orchestrated all in Goose, Gemini picked four options that fit this message better than a generic list would, and it followed the skill’s advice to define what yes and no mean for the deadline question. Jev ascribed certainty (or at least a confidence of 1) on cloud migration and 0.98 on the deadline. Gemini then interpreted the June lease and the server move as its evidence, a rationale that Jev never gives because it returns only typed values.

To be clear, Gemini described the answer with “100% confidence,” but TypeSafe’s guidance says confidence only summarizes how concentrated the probabilities are. It does not tell you whether the answer is correct.

A Choice maps onto a switch statement, and a Noul maps onto an if. Your code can move this message to the top of the inbox before anyone reads it, and Gemini can explain the decision to anyone who asks.

The pairing also keeps every layer replaceable. Goose let me choose Gemini for the conversation, and it would let me choose Claude or a local model just as easily. The skill follows the open Agent Skills standard, and Jev sits behind a single API endpoint. This is the beauty of standards like Goose, Agent Skills, and even old-school standards like HTTP. You can swap the LLM and decision model (for example, to replace Jev with the open source Laya decision model) without changing your workflow.

The pairing has one cost to watch. For a single question, the $0.07 that Gemini cost dwarfs the fraction of a cent that TypeSafe charges for Jev. Gemini has to read the skill, the reference, and the whole conversation to plan each step. The fix is to give loops to code. When a skill ships a script that sends every item in a batch to Jev, the LLM plans the run and explains the results while Jev makes every decision in between. Your LLM bill then tracks conversations, and your Jev bill tracks decisions.

This is your regular reminder that old-fashioned code will always be your best option.

A Real Business Case for Vidya: Triage for Sources Sought Notices

Vidya serves government customers as well as commercial businesses, and Sources Sought is a key marketing strategy for us.

Before they write a solicitation, federal agencies publish Sources Sought notices on SAM.gov. A Sources Sought notice is market research. The contracting officer uses the responses to learn which businesses can do the work and whether enough capable small businesses exist to set the contract aside for them. A response never wins a contract on its own, but it can shape the requirement and the set-aside decision before the competition begins. For a small business like Vidya, those notices are some of the most valuable ways to get noticed.

My finding-sources-sought skill searches SAM.gov through the official Opportunities API with one call for each of Vidya’s six North American Industry Classification System (NAICS) codes. It drops notices that announce a sole-source award and notices that require a certification Vidya lacks. The last step checks each surviving notice against our capabilities in the Vidya knowledge base, and it is the hardest step to get right.

NAICS codes are broad, so SAM.gov has returned Cisco SmartNet hardware maintenance and Dell tower computer purchases under Vidya’s software codes. The skill runs a hybrid search over the knowledge base that blends keyword matching with embeddings. Words like maintenance, support, and training appear in nearly every federal notice and in nearly every one of Vidya’s past responses. That shared vocabulary makes it hard for any search to separate a real match from a coincidence.

Just for fun, I ran the skill’s knowledge base search today on two notices it had already processed. The best matches for a contract to maintain a building controls system, which is not our lane at all, and a software modernization effort, which is our lane exactly, scored within 5 percent of each other. They should not be anywhere near each other in the rankings. I could try to update the way the Vidya knowledge base works or tweak the language of the skill, but those solutions treat the symptom rather than cure the disease, which is that LLMs are not built for decision making.

I updated finding-sources-sought to use Jev instead of LLM judgment to identify opportunities that truly align with our core capabilities.

As the skill pores through the array of opportunities in the response from the SAM API, it can hand each one to Jev for evaluation. Gemini orchestrates everything. Deterministic rules such as the response deadline and the set-aside code stay in code. Jev makes the decision on each opportunity, and Gemini turns the combined results into a judgment for me in plain English.

Each notice gets one Jev request with three questions. Its state holds the notice and the seven service lines in Vidya’s capability statement from the knowledge base. A Choice names the kind of work the notice requests. A Noul asks whether the agency has already decided on a sole-source award. A Score rates how well the notice fits Vidya’s services, and its levels spell out the difference between shared subject matter and shared vocabulary.

The Choice takes its options from the knowledge base, one for each kind of work Vidya pursues: modernization, architecture, AI, user experience, data, training, cybersecurity, and DevSecOps. TypeSafe advises an other option when the list might not cover every input, but I left it out because the Score decides whether a notice fits. Jev cannot pick an option you leave out, which is what the type safety is for, so it labels every notice with one of those eight.

{
  "model": "jev-latest",
  "state": {
    "notice": {
      "title": "Automation and Modernization",
      "description_excerpt": "..."
    },
    "vidya_capabilities": [
      "Architecture modernization to use industry-leading technologies and get the most out of valuable legacy data.",
      "Software development at cloud scale to make services available to anyone around the world.",
      "Elegant, accessible web development to showcase the brand with memorable user experiences.",
      "Machine learning and AI to help everyone do things faster.",
      "Cybersecurity and Zero Trust to protect everyone's assets.",
      "Engineering automation through DevSecOps to deliver fast at high quality.",
      "Engineering training courses to help anyone change the world with technology."
    ]
  },
  "questions": {
    "work_type": {
      "type": "choice",
      "instructions": "What kind of work does `notice` ask a contractor to perform?",
      "criteria": {
        "modernization": "Modernize legacy systems and integrate them with modern technologies",
        "architecture": "Design and build software, APIs, and platforms at cloud scale",
        "ai": "Build machine learning and AI solutions, including agents, that help people do things faster",
        "user_experience": "Design and build accessible websites and web applications with memorable user experiences",
        "data": "Engineer data pipelines, platforms, and analytics that get the most out of valuable data",
        "training": "Design and deliver engineering training courses that help people change the world with technology",
        "cybersecurity": "Protect systems and data with cybersecurity and Zero Trust",
        "devsecops": "Automate engineering through DevSecOps to deliver software fast at high quality"
      }
    },
    "sole_source_decided": {
      "type": "noul",
      "instructions": "Does `notice` say the agency has already decided to award the work to a specific contractor?",
      "criteria": {
        "true": "States an intent to award a sole-source contract to a named or identified contractor",
        "false": "Has no sole-source language, or only says the agency may consider a sole-source award"
      }
    },
    "capability_fit": {
      "type": "score",
      "instructions": "How well does the work that `notice` requests fit the services in `vidya_capabilities`?",
      "criteria": [
        "None of the services covers this kind of work",
        "A service shares words with the notice, such as maintenance, support, or training, but covers a different kind of work",
        "A service covers related work, but the notice centers on something else",
        "A service covers exactly this kind of work"
      ]
    }
  }
}

These are the results.

// Maintenance and Emergency Services
"work_type": {
  "type": "choice",
  "choice": "modernization",
  "confidence": 0.54,
  "probabilities": { "devsecops": 0.03, "cybersecurity": 0.05, "user_experience": 0.04, "architecture": 0.23, "data": 0.0, "training": 0.03, "modernization": 0.61, "ai": 0.01 }
},
"sole_source_decided": { "type": "noul", "noul": 0.03 },
"capability_fit": {
  "type": "score",
  "score": 1.12,
  "confidence": 0.57,
  "probabilities": { "0": 0.15, "1": 0.63, "2": 0.17, "3": 0.05 }
}

// Automation and Modernization
"work_type": {
  "type": "choice",
  "choice": "modernization",
  "confidence": 0.99,
  "probabilities": { "devsecops": 0.01, "ai": 0.0, "architecture": 0.0, "user_experience": 0.0, "modernization": 0.99, "data": 0.0, "cybersecurity": 0.0, "training": 0.0 }
},
"sole_source_decided": { "type": "noul", "noul": 0.03 },
"capability_fit": {
  "type": "score",
  "score": 2.89,
  "confidence": 0.89,
  "probabilities": { "0": 0.0, "1": 0.0, "2": 0.09, "3": 0.91 }
}

Jev put the modernization work at 2.89 out of 3 and the maintenance work at 1.12. That gap is what the filter in my skill needs to be useful.

What you do with the result is up to you. TypeSafe recommends keeping every question and threshold in one file so a person can review them, and the policy for this skill fits in a few lines of TypeScript.

type WorkType =
  | "modernization"
  | "architecture"
  | "ai"
  | "user_experience"
  | "data"
  | "training"
  | "cybersecurity"
  | "devsecops";

type NoticeAnswers = Readonly<{
  work_type: Readonly<{ choice: WorkType; confidence: number }>;
  sole_source_decided: Readonly<{ noul: number }>;
  capability_fit: Readonly<{ score: number; confidence: number }>;
}>;

type Triage =
  | Readonly<{ kind: "pursue"; work: WorkType; fit: number }>
  | Readonly<{ kind: "drop"; reason: string }>
  | Readonly<{ kind: "review"; reason: string }>;

// Starting points to calibrate against past go and no-go decisions
const SOLE_SOURCE_CUTOFF = 0.9;
const MIN_CONFIDENCE = 0.8;
const MIN_FIT = 1.5; // Score levels run from 0 to 3

export const triage = ({ work_type, sole_source_decided, capability_fit }: NoticeAnswers): Triage => {
  if (sole_source_decided.noul >= SOLE_SOURCE_CUTOFF) return { kind: "drop", reason: "sole source" };
  if (capability_fit.score < MIN_FIT) return { kind: "drop", reason: "weak capability fit" };
  if (work_type.confidence < MIN_CONFIDENCE) return { kind: "review", reason: "uncertain work type" };
  return { kind: "pursue", work: work_type.choice, fit: capability_fit.score };
};

That policy puts the two notices far apart. The maintenance work drops out because its fit of 1.12 falls below the 1.5 cutoff. The modernization work notice passes every check and enters the report as a fit of 2.89 out of 3.

Gemini’s judgment then names the opportunities worth a response and carries the probabilities behind each recommendation. I can see why the skill kept or dropped each notice without rereading an LLM’s reasoning, and your team gets the same audit trail for any decision it automates this way.

Start with One Decision

Most of your work is making decisions, and the decisions your agents make have a short list of acceptable answers. Keep your LLM for the conversation, and pick one of those decisions, such as ticket routing or lead triage, to rewrite as a single Choice or Noul question. Then send the same 50 real cases to Jev and to the frontier model you use today. Within a week, you will know what your agents pay for each snap judgment and whether they need to keep paying it.

Want to transform your business?Get in touch today!

Partner with Vidya to modernize your technology, empower your teams, and accelerate your mission.