{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "@id": "https://www.vidyasource.com/blog/making-decisions-with-jev-and-goose/",
  "url": "https://www.vidyasource.com/blog/making-decisions-with-jev-and-goose/",
  "mainEntityOfPage": "https://www.vidyasource.com/blog/making-decisions-with-jev-and-goose/",
  "headline": "Making Decisions with Jev and Goose",
  "description": "Pair an LLM in Goose with TypeSafe's Jev, and your AI agents talk in plain English while they make fast, calibrated decisions for a fraction of a cent.",
  "datePublished": "2026-09-30T00:00:00.000Z",
  "author": {
    "@type": "Person",
    "name": "Neil Chaudhuri",
    "jobTitle": "President",
    "url": "https://www.linkedin.com/in/neil-chaudhuri/",
    "sameAs": "https://www.linkedin.com/in/neil-chaudhuri/"
  },
  "publisher": {
    "@id": "https://www.vidyasource.com/#organization"
  },
  "image": "https://www.vidyasource.com/img/blog/making-decisions-with-jev-and-goose.webp",
  "keywords": [
    "AI",
    "Agentic AI",
    "Agents",
    "Agent Skills",
    "Machine Learning",
    "Open Source",
    "Government",
    "Security",
    "Architecture",
    "TypeScript"
  ],
  "articleBody": "Most of the work your business runs on comes down to decisions. Your teams route support tickets, approve or flag invoices, and triage security alerts.\nYour AI agents spend their days on the same kind of work, and nearly every one of those tasks has a short list of acceptable answers.\nNone of them needs a model that can write poems from the perspective of Voldemort in Harry Potter.\n\nMany teams send those decisions to the same frontier LLM that drafts their emails. An LLM writes its answer one token at a time,\nso even a one-word label arrives as generated text that your code must parse and validate before it can act. You wait for and pay for every one of those tokens.\n\nA decision model answers the question directly. Jev, the first decision model from TypeSafe, takes your application's state along with a set of typed\nquestions and returns typed answers with probabilities in a fraction of a second. You can run Jev from Goose, the open source agent that Block contributed to the\n[Agentic AI Foundation](https://www.vidyasource.com/blog/agentic-ai-foundation-ambassador/) (AAIF), with an LLM that orchestrates the work flow and interacts with you through natural language.\nI believe that pairing beats either model on its own. The LLM speaks\nnatural language with the people who use the agent, and the decision model makes cheap, fast, calibrated calls for the software behind it.\n\n## The Power of a Decision Model\n\nA recent [Hugging Face guide](https://huggingface.co/blog/sora-2/jev-ai-vs-llms-when-should-you-use-a-decision-mode) to choosing between Jev and an LLM\nmakes the distinction concrete. An LLM generates content for people to read, and even its structured output remains generated text that your code must\ncheck for missing fields and formatting drift. A decision model starts from an answer space you define and returns a yes-or-no\nprobability, choice, or\na score for software to act on. The guide offers a simple test. If you can list the possible answers in advance and your code will route, rank, or\nblock on the result, use a decision model.\n\nTypeSafe calls the building blocks of that answer space [primitives](https://docs.typesafe.ai/primitives), and each one pairs a question type with a typed answer:\n\n- **Noul** answers a yes-or-no question with the probability of yes. \"Is the customer satisfied?\"\n- **Choice** picks one option from a set you define and returns the top option, the probability of every option, and a confidence value. \"Is the customer most likely to buy drinks, appetizers, meals, or desserts?\"\n- **Score** places the state on an ordered scale of levels you describe and returns a probability-weighted position on that scale. \"On a scale from Kid Rock to Taylor Swift, how much did the customer like this song?\"\n\nOne request can ask several questions about the same state, and Jev answers each one independently under an identifier you choose. Your code can ask for\na department, an urgency level, and a refund flag in a single call and then tune the threshold for each answer without touching the others.\n\nTypeSafe calls Jev a [System One model](https://docs.typesafe.ai/concepts/system-one), a name that borrows from the fast, automatic System 1 in Daniel Kahneman's\n[account of human judgment](https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow). TypeSafe trains these models\nwith a method it calls Reinforcement Learning for Calibrated Decisions (RLCD) to provide the kind of judgment an expert makes in about a second. The [home page](https://typesafe.ai) argues that the Reinforcement\nLearning from Human Feedback (RLHF) behind chat models makes them overconfident. Calibration means the probabilities track real outcomes across many\npredictions. TypeSafe's documentation also cautions that calibration cannot guarantee any single answer, so your code should weigh each probability as\nevidence before it acts.\n\n## The Need for Speed\n\nOver the last few years, we have delegated decisions to frontier models with type safety optional although I am personally [relentless about type safety](https://www.vidyasource.com/blog/python-is-not-language-of-ai/).\nIt was the best we could do even if we know deep down the judgments LLMs give us and the probabilities they may ascribe to their judgments are more or less\nconfident fiction.\n\nThe Hugging Face guide suggests a decision model whenever a workflow needs many small judgments, routing by probability, or a decision interface that holds steady\nat high volume. I completely agree because decision models are built for that while LLMs are not.\n\nTypeSafe's [workflow evaluations](https://evals.typesafe.ai) put numbers on that choice. The evaluations cover 705 cases across four automation workflows,\nfrom invoice fraud triage to security alert triage, and every model runs at its provider's default reasoning setting. Across all four workflows, Jev matched\nthe 67.8 percent average accuracy of Claude Sonnet 5 for $0.0004 per case instead of $0.1174. Jev also finished each case in 0.4 seconds against\n78.1 seconds for Claude Sonnet 5. At those prices, your agent can run more than 250 cases through Jev for the cost of one case on Claude Sonnet 5.\n\nThe same evaluations make a second point that matters even more. Every model that TypeSafe tested in both modes scored higher as a structured workflow of\nsmall decisions than as one standalone prompt. Claude Haiku 4.5 alone jumped from 18.1 percent to 53.6 percent. Your agents gain accuracy as soon as you\nbreak their work into decisions, and a decision model makes each of those decisions cheap.\n\nTypeSafe publishes the\n[harness on GitHub](https://github.com/typesafe-ai/WorkflowEvals) and the\n[datasets on Hugging Face](https://huggingface.co/collections/typesafe/workflowevals-6abaf8dcb1e283d9e3d57b6b), so your team can rerun the comparison on\nyour own cases before committing. My own measurements are consistent with theirs. The two requests behind my sources sought example below made the round trip from my Mac to TypeSafe in 169\nand 184 milliseconds, network included. At that speed, a check before each tool call adds less than a fifth of a second to your agent's work, which I\nconsider a bargain for a calibrated answer.\n\n## Teaching Goose to Use Jev\n\nI cannot stress enough how much AI demands interoperability so we can enjoy freedom of choice with our tools, and [Goose](https://goose-docs.ai/) shines here.\nIt also supports [Agent Skills](https://goose-docs.ai/docs/guides/context-engineering/using-skills), the open standard for packaging instructions, domain expertise, and\nresources that an agent loads on demand, and it discovers global skills in `~/.agents/skills`. A skill you write for one agent carries over to Goose. This is the\nbeauty of standards, which I am also [relentless](https://www.vidyasource.com/blog/agentic-ai-foundation-ambassador/) about.\n\nTypeSafe publishes an [agent skill](https://docs.typesafe.ai/agent-skill) that teaches a coding agent the three question types, their architectural\npatterns, and their practices for evaluating results. Any agent that supports the standard can install it with one command.\n\n```bash\nnpx skills add typesafe-ai/skills --skill typesafe-ai\n```\n\nI made two changes before I let Goose use it.\n\n### Bundled References\n\nTypeSafe's skill points the agent to its `llms.txt` documentation index and has it fetch pages such as the API\nreference and the guide for each primitive from their site. I replaced those links with a `references` folder inside the skill, a layout\nfor [progressive disclosure](https://www.vidyasource.com/blog/put-your-opinions-to-work-with-agents-md-and-keep-them-safe/) that the\n[Agent Skills specification](https://agentskills.io/specification) also recommends for documentation. The folder holds a copy of TypeSafe's API reference and\na condensed guide to writing questions. The skill tells Goose to read those files and to fetch nothing without asking me first. Goose now works\nfrom documentation I reviewed instead of whatever a web page somewhere says at the moment it loads.\n\nThat difference matters for security. Every page an agent fetches can carry a [prompt injection](https://genai.owasp.org/llmrisk/llm01-prompt-injection/),\nthe first risk in the OWASP Top 10 for LLM Applications. I have\n[warned before](https://www.vidyasource.com/blog/put-your-opinions-to-work-with-agents-md-and-keep-them-safe/) that borrowed agent content is a vulnerability. The bundled folder also saves\na network round trip on every task, and in some of my environments it saves an approval prompt too.\n\nBundled references do offer one risk. TypeSafe's own troubleshooting guide warns that a stale skill can lead an agent to invent request or response\nfields, and my copy starts to age the day TypeSafe changes its API. Every Jev response names the model that produced it, `jev-1.13.0` as I write this,\nso a change in that field tells you it is time to compare the folder against TypeSafe's documentation. This is your regular reminder\nthat we still have to wrestle with breaking changes in the age of AI.\n\n### A Key from the Keychain\n\nTypeSafe's quickstart expects the API key in the `TYPESAFE_API_KEY` environment variable. An environment variable travels to every process the shell\nstarts, and an agent that prints its environment can copy the key straight into a transcript. My version of the skill tells Goose to read the key from the\nmacOS keychain through the [Computer Controller extension](https://goose-docs.ai/docs/mcp/computer-controller-mcp), which ships with Goose and automates\nmacOS applications and system tasks. The keychain encrypts the key at rest, and the skill has Goose read the key and send the request in one command so\nthe value never lands in the transcript.\n\nStoring the key is easy. With `-w` at the end, `security` prompts for the value and keeps it out of your shell history.\n\n```bash\nsecurity add-generic-password -a \"$USER\" -s \"TYPESAFE_API_KEY\" -w\n```\n\nThe call reads the key and uses it in the same step. Piping the header into `curl -K -` also keeps the key out of the process list.\n\n```bash\nKEY=\"$(security find-generic-password -a \"$USER\" -s \"TYPESAFE_API_KEY\" -w)\"\nprintf 'header = \"Authorization: Bearer %s\"\\n' \"$KEY\" |\n  curl -sS -K - https://api.typesafe.ai/v1/systemone \\\n    -H \"Content-Type: application/json\" \\\n    --data-binary @request.json\nunset KEY\n```\n\nThe same pattern works with any secret store your platform offers, and it keeps every credential your agents use out of their transcripts.\n\n## A Trivial First Decision\n\nA trivial example shows how the two models split the work. I asked Goose, running Gemini 3.1 Pro, to judge the sentiment of one\nsentence about my favorite baseball team.\n\n```markdown\nUsing the typesafe skill tell me if this is a positive sentiment: \"My favorite\nbaseball team is in the top 10 in team OPS in Major League Baseball.\"\n```\n\nGoose loaded the `typesafe-ai` skill and its `references/api.md` file, ran one `curl` command against the TypeSafe API, and returned this response:\n\n```markdown\nUsing the `typesafe-ai` skill, I evaluated the statement with the `jev-latest` model.\n\n**Yes, this is a positive sentiment.**\n\nThe model returned a `noul` (probability) of **0.96**, indicating a 96% certainty that the statement expresses a positive, happy, or proud sentiment. Being in the top 10 for team OPS (On-base Plus Slugging) is a strong offensive statistic in baseball, making it a positive statement to make about a favorite team!\n```\n\nThe whole exchange finished within the same minute,\nand Goose reported $0.07 of Gemini usage for it. Gemini also explained that a Top 10 rank in OPS marks a strong offense, which is true. That is context\nthat Jev may not know and would never tell me anyway because it returns only typed values.\n\nThis might be a silly example, but it gives you a sense of how decision models and LLMs complement each other so well.\n\n## Gemini Talks, Jev Decides\n\nThat exchange gave each model the job it does best. Gemini handled every step that needed language. It understood a request in natural language.\nIt then wrote a valid API request from the bundled reference and explained Jev's answer back to me in natural language. Jev handled the one step that\nneeded a judgment. The result is a cheap, fast decision that an LLM orchestrates and explains in a manner I can easily understand.\n\nThe typed answer matters as much as the explanation because it lets the same judgment move from a conversation into code. Gemini translates natural language into\nthe types Jev needs and then interprets the types in Jev's response to the language I need.\n\nHere is a more interesting example.\n\n```markdown\nUsing the typesafe skill, tell me what the sender of this contact form message wants from\nVidya and whether the message names a deadline: \"Hi Neil, our team runs a 15-year-old Java\napplication on servers we have to vacate next year, and nobody left on the team knows how it\nworks. Could Vidya help us move it to the cloud before our data center lease ends in June?\"\n```\n\nGemini designed the questions itself and turned that prompt into this request.\n\n```json\n{\n  \"model\": \"jev-latest\",\n  \"state\": \"Hi Neil, our team runs a 15-year-old Java application on servers we have to vacate next year, and nobody left on the team knows how it works. Could Vidya help us move it to the cloud before our data center lease ends in June?\",\n  \"questions\": {\n    \"service_requested\": {\n      \"type\": \"choice\",\n      \"instructions\": \"What primary service is the sender requesting from Vidya?\",\n      \"criteria\": {\n        \"cloud_migration\": \"Moving an existing legacy application or system to cloud infrastructure.\",\n        \"new_development\": \"Building a brand new software application from scratch.\",\n        \"team_training\": \"Training or upskilling the existing team on new technologies.\",\n        \"unclear\": \"The requested service is not clear or is something else.\"\n      }\n    },\n    \"has_deadline\": {\n      \"type\": \"noul\",\n      \"instructions\": \"Does the message explicitly name a deadline or time constraint?\",\n      \"criteria\": {\n        \"true\": \"The sender states a specific month, timeline, or event by which the work must be completed.\",\n        \"false\": \"The sender does not state any time constraints or deadlines.\"\n      }\n    }\n  }\n}\n```\n\nJev returned typed answers under the identifiers that Gemini chose.\n\n```json\n{\n  \"model\": \"jev-1.13.0\",\n  \"answers\": {\n    \"service_requested\": {\n      \"type\": \"choice\",\n      \"choice\": \"cloud_migration\",\n      \"confidence\": 1.0,\n      \"probabilities\": {\n        \"unclear\": 0.0,\n        \"new_development\": 0.0,\n        \"cloud_migration\": 1.0,\n        \"team_training\": 0.0\n      }\n    },\n    \"has_deadline\": {\n      \"type\": \"noul\",\n      \"noul\": 0.98\n    }\n  },\n  \"usage\": {\n    \"input_tokens\": 502,\n    \"output_tokens\": 78\n  }\n}\n```\n\nGemini then translated those answers back into natural language.\n\n```markdown\nUsing the `typesafe-ai` skill, here is the structured evaluation of the contact form message:\n\n- **What they want:** The model determined with **100% confidence** (`choice: \"cloud_migration\"`) that the sender wants help moving their legacy application to cloud infrastructure.\n- **Is there a deadline: Yes**. The model gave a **0.98** probability (`noul: 0.98`) that the message explicitly names a deadline or time constraint. (Specifically, the text mentions having to vacate servers next year and the lease ending in June).\n```\n\nOrchestrated all in Goose, Gemini picked four options that fit this message better than a generic list would, and it followed the skill's advice to define what yes and no mean for\nthe deadline question. Jev ascribed certainty (or at least a confidence of 1) on cloud migration and 0.98 on the deadline. Gemini then interpreted the June lease and the server move as\nits evidence, a rationale that Jev never gives because it returns only typed values.\n\nTo be clear, Gemini described the answer with\n\"100% confidence,\" but TypeSafe's guidance says confidence only summarizes how concentrated the probabilities are. It does not tell you whether the answer\nis correct.\n\nA Choice maps onto a `switch` statement, and a Noul maps onto an `if`. Your code can move this message to the top of the inbox before anyone reads it, and\nGemini can explain the decision to anyone who asks.\n\nThe pairing also keeps every layer replaceable. Goose let me choose Gemini for the conversation, and it would let me choose Claude or a local model just\nas easily. The skill follows the open Agent Skills standard, and Jev sits behind a single API endpoint. This is the beauty of standards like Goose, Agent Skills, and\neven old-school standards like HTTP. You can swap the LLM and decision model (for example, to replace Jev with the open source\n[Laya decision model](https://huggingface.co/blog/sora-2/laya-ai-model-how-it-works-run-it-locally-and-eval)) without changing your workflow.\n\nThe pairing has one cost to watch. For a single question, the $0.07 that Gemini cost dwarfs the fraction of a cent that TypeSafe charges for Jev. Gemini\nhas to read the skill, the reference, and the whole conversation to plan each step. The fix is to give loops to code. When a skill ships a script that sends\nevery item in a batch to Jev, the LLM plans the run and explains the results while Jev makes every decision in between. Your LLM bill then tracks\nconversations, and your Jev bill tracks decisions.\n\nThis is your regular reminder that old-fashioned code will always be your best option.\n\n## A Real Business Case for Vidya: Triage for Sources Sought Notices\n\nVidya serves government customers as well as commercial businesses, and Sources Sought is a key marketing strategy for us.\n\nBefore they write a solicitation, federal agencies publish Sources Sought notices on [SAM.gov](https://sam.gov).\nA Sources Sought notice is market research. The contracting officer uses the responses to learn which businesses can do the work and whether\nenough capable small businesses exist to set the contract aside for them. A response never wins a contract on its own, but it can shape the requirement\nand the set-aside decision before the competition begins. For a small business like Vidya, those notices are some of the most valuable ways to get noticed.\n\nMy `finding-sources-sought` skill searches SAM.gov through the official [Opportunities API](https://open.gsa.gov/api/get-opportunities-public-api/) with one call for\neach of Vidya's six North American Industry Classification System (NAICS) codes. It drops notices that announce a sole-source award and notices that require\na certification Vidya lacks. The last step checks each surviving notice against our capabilities in the Vidya knowledge base, and it is the hardest step to\nget right.\n\nNAICS codes are broad, so SAM.gov has returned Cisco SmartNet hardware maintenance and Dell tower computer purchases under Vidya's software codes. The skill\nruns a hybrid search over the knowledge base that blends keyword matching with embeddings. Words like maintenance, support, and training appear in nearly\nevery federal notice and in nearly every one of Vidya's past responses. That shared vocabulary makes it hard for any search to separate a real match from\na coincidence.\n\nJust for fun, I ran the skill's knowledge base search today on two notices it had already processed. The best matches for a contract to\nmaintain a building controls system, which is not our lane at all, and a software modernization effort, which is our lane exactly, scored within 5\npercent of each other. They should not be anywhere near each other in the rankings. I could try to update the way the Vidya knowledge base works or tweak the language of the skill, but those\nsolutions treat the symptom rather than cure the disease, which is that LLMs are not built for decision making.\n\nI updated `finding-sources-sought` to use Jev instead of LLM judgment to identify opportunities that truly align with our core capabilities.\n\nAs the skill pores through the array of opportunities in the response from the SAM API, it can hand each one to Jev for evaluation.\nGemini orchestrates everything. Deterministic rules such as the response deadline and the set-aside code stay in code. Jev makes the decision on each\nopportunity, and Gemini turns the combined results into a judgment for me in plain English.\n\nEach notice gets one Jev request with three questions. Its state holds the notice and the seven service lines in Vidya's capability statement from the\nknowledge base. A `Choice` names the kind of work the notice requests. A `Noul` asks whether the agency has already decided on a sole-source award. A\n`Score` rates how well the notice fits Vidya's services, and its levels spell out the difference between shared subject matter and shared vocabulary.\n\nThe `Choice` takes its options from the knowledge base, one for each kind of work Vidya pursues: modernization, architecture, AI, user\nexperience, data, training, cybersecurity, and DevSecOps. [TypeSafe advises](https://docs.typesafe.ai/primitives/choice) an `other` option when the list\nmight not cover every input, but I left it out because the `Score` decides whether a notice fits. Jev cannot pick an option you leave out, which is what the\ntype safety is for, so it labels every notice with one of those eight.\n```json\n{\n  \"model\": \"jev-latest\",\n  \"state\": {\n    \"notice\": {\n      \"title\": \"Automation and Modernization\",\n      \"description_excerpt\": \"...\"\n    },\n    \"vidya_capabilities\": [\n      \"Architecture modernization to use industry-leading technologies and get the most out of valuable legacy data.\",\n      \"Software development at cloud scale to make services available to anyone around the world.\",\n      \"Elegant, accessible web development to showcase the brand with memorable user experiences.\",\n      \"Machine learning and AI to help everyone do things faster.\",\n      \"Cybersecurity and Zero Trust to protect everyone's assets.\",\n      \"Engineering automation through DevSecOps to deliver fast at high quality.\",\n      \"Engineering training courses to help anyone change the world with technology.\"\n    ]\n  },\n  \"questions\": {\n    \"work_type\": {\n      \"type\": \"choice\",\n      \"instructions\": \"What kind of work does `notice` ask a contractor to perform?\",\n      \"criteria\": {\n        \"modernization\": \"Modernize legacy systems and integrate them with modern technologies\",\n        \"architecture\": \"Design and build software, APIs, and platforms at cloud scale\",\n        \"ai\": \"Build machine learning and AI solutions, including agents, that help people do things faster\",\n        \"user_experience\": \"Design and build accessible websites and web applications with memorable user experiences\",\n        \"data\": \"Engineer data pipelines, platforms, and analytics that get the most out of valuable data\",\n        \"training\": \"Design and deliver engineering training courses that help people change the world with technology\",\n        \"cybersecurity\": \"Protect systems and data with cybersecurity and Zero Trust\",\n        \"devsecops\": \"Automate engineering through DevSecOps to deliver software fast at high quality\"\n      }\n    },\n    \"sole_source_decided\": {\n      \"type\": \"noul\",\n      \"instructions\": \"Does `notice` say the agency has already decided to award the work to a specific contractor?\",\n      \"criteria\": {\n        \"true\": \"States an intent to award a sole-source contract to a named or identified contractor\",\n        \"false\": \"Has no sole-source language, or only says the agency may consider a sole-source award\"\n      }\n    },\n    \"capability_fit\": {\n      \"type\": \"score\",\n      \"instructions\": \"How well does the work that `notice` requests fit the services in `vidya_capabilities`?\",\n      \"criteria\": [\n        \"None of the services covers this kind of work\",\n        \"A service shares words with the notice, such as maintenance, support, or training, but covers a different kind of work\",\n        \"A service covers related work, but the notice centers on something else\",\n        \"A service covers exactly this kind of work\"\n      ]\n    }\n  }\n}\n```\n\nThese are the results.\n\n```jsonc\n// Maintenance and Emergency Services\n\"work_type\": {\n  \"type\": \"choice\",\n  \"choice\": \"modernization\",\n  \"confidence\": 0.54,\n  \"probabilities\": { \"devsecops\": 0.03, \"cybersecurity\": 0.05, \"user_experience\": 0.04, \"architecture\": 0.23, \"data\": 0.0, \"training\": 0.03, \"modernization\": 0.61, \"ai\": 0.01 }\n},\n\"sole_source_decided\": { \"type\": \"noul\", \"noul\": 0.03 },\n\"capability_fit\": {\n  \"type\": \"score\",\n  \"score\": 1.12,\n  \"confidence\": 0.57,\n  \"probabilities\": { \"0\": 0.15, \"1\": 0.63, \"2\": 0.17, \"3\": 0.05 }\n}\n\n// Automation and Modernization\n\"work_type\": {\n  \"type\": \"choice\",\n  \"choice\": \"modernization\",\n  \"confidence\": 0.99,\n  \"probabilities\": { \"devsecops\": 0.01, \"ai\": 0.0, \"architecture\": 0.0, \"user_experience\": 0.0, \"modernization\": 0.99, \"data\": 0.0, \"cybersecurity\": 0.0, \"training\": 0.0 }\n},\n\"sole_source_decided\": { \"type\": \"noul\", \"noul\": 0.03 },\n\"capability_fit\": {\n  \"type\": \"score\",\n  \"score\": 2.89,\n  \"confidence\": 0.89,\n  \"probabilities\": { \"0\": 0.0, \"1\": 0.0, \"2\": 0.09, \"3\": 0.91 }\n}\n```\n\nJev put the modernization work at 2.89 out of 3 and the maintenance work at 1.12. That gap is what the filter in my skill needs to be useful.\n\nWhat you do with the result is up to you. TypeSafe recommends keeping every question and threshold in one file so a person can review them, and the policy for this skill fits in a few lines of\nTypeScript.\n\n```typescript\ntype WorkType =\n  | \"modernization\"\n  | \"architecture\"\n  | \"ai\"\n  | \"user_experience\"\n  | \"data\"\n  | \"training\"\n  | \"cybersecurity\"\n  | \"devsecops\";\n\ntype NoticeAnswers = Readonly<{\n  work_type: Readonly<{ choice: WorkType; confidence: number }>;\n  sole_source_decided: Readonly<{ noul: number }>;\n  capability_fit: Readonly<{ score: number; confidence: number }>;\n}>;\n\ntype Triage =\n  | Readonly<{ kind: \"pursue\"; work: WorkType; fit: number }>\n  | Readonly<{ kind: \"drop\"; reason: string }>\n  | Readonly<{ kind: \"review\"; reason: string }>;\n\n// Starting points to calibrate against past go and no-go decisions\nconst SOLE_SOURCE_CUTOFF = 0.9;\nconst MIN_CONFIDENCE = 0.8;\nconst MIN_FIT = 1.5; // Score levels run from 0 to 3\n\nexport const triage = ({ work_type, sole_source_decided, capability_fit }: NoticeAnswers): Triage => {\n  if (sole_source_decided.noul >= SOLE_SOURCE_CUTOFF) return { kind: \"drop\", reason: \"sole source\" };\n  if (capability_fit.score < MIN_FIT) return { kind: \"drop\", reason: \"weak capability fit\" };\n  if (work_type.confidence < MIN_CONFIDENCE) return { kind: \"review\", reason: \"uncertain work type\" };\n  return { kind: \"pursue\", work: work_type.choice, fit: capability_fit.score };\n};\n```\n\nThat policy puts the two notices far apart. The maintenance work drops out because its fit of 1.12 falls below the 1.5 cutoff. The modernization work\nnotice passes every check and enters the report as a fit of 2.89 out of 3.\n\nGemini's judgment then names the opportunities worth a response and carries the probabilities behind each recommendation. I can see why the skill kept or dropped each notice without\nrereading an LLM's reasoning, and your team gets the same audit trail for any decision it automates this way.\n\n## Start with One Decision\n\nMost of your work is making decisions, and the decisions your agents make have a short list of acceptable answers. Keep your LLM for the conversation, and pick one of those decisions,\nsuch as ticket routing or lead triage, to rewrite as a single Choice or Noul question. Then send the same 50 real cases to Jev and to the frontier model you use today. Within a week, you will know what your\nagents pay for each snap judgment and whether they need to keep paying it."
}