TypeSafe AI shipped Jev on September 15. Eleven days later, we co-hosted the first hackathon built around it, with GMI Cloud inference underneath the projects
October 01, 2026

Jev is the first model from TypeSafe AI, a company started in 2024 by Diogo Almeida, who spent four years at OpenAI on RLHF, InstructGPT and ChatGPT. It went into early access on September 15, 2026, and within days people were calling it one of the fastest adopted launches in AI. On Saturday, September 26, GMI Cloud co-hosted JEVATHON with TypeSafe AI and The AI Collective, the first community hackathon built around it, at the CodeRabbit office in San Francisco, with the Bay Bridge in the window behind the demo stage.
500+ people RSVP'd on Luma over five days, and 55 teams shipped a working project in three hours of build time. Hendrik Krack from CodeRabbit called it on camera: "It's the first Jev hackathon that was ever hosted in the world."

GMI Cloud's part was the inference layer. Every team got credits on our MaaS inference engine, so the LLM, image, video and voice calls behind each project ran on GMI Cloud endpoints. I ran the morning workshop on how to combine Jev with those models, and I sat on the judging panel in the afternoon.
Jev is a different kind of model, and the workshop started there. TypeSafe calls it a System One model, after Kahneman's fast, intuitive System 1 thinking. You send it state (a string, a JSON blob, a transcript) plus a typed question, and it returns a typed answer with a probability attached. There are three primitives: Choice picks one of the options you defined, Score rates against an ordered scale, and a yes/no check. Because the output schema is fixed in advance, the answer is always a valid value your code can branch on. TypeSafe quotes 70 to 500 ms per call and input at $0.042 per million tokens with free output.

Allie Laabs from TypeSafe had the shortest version: Jev is a smart if statement, an intelligence-infused logic gate you can put inside the software loop because it is fast and cheap enough to call on every tick. The name comes from Jevons Paradox, the idea that when the price of something drops, the uses for it multiply.
What Jev leaves out is generation. It will tell you which of five actions to take, or whether a support ticket is urgent, or how a frame scores against a rubric. It will write you zero sentences. That is the gap GMI Cloud filled. The pattern I taught was simple: every branch point in the agent goes to Jev, every step that has to produce text, an image, a video or a voice goes to a model on GMI Cloud MaaS, and the two talk through the same OpenAI-compatible client.


Teams had the full GMI Cloud MaaS catalog on the same credits: open text models like DeepSeek, Qwen and GLM, closed ones like Claude and GPT, plus image, video and voice models, all behind one OpenAI-compatible endpoint.
The hands-on part covered three things. First, a Jev call with a Choice primitive wired to a GMI Cloud chat completion, so the audience saw the decision come back and the generated text follow it. Second, swapping the model behind a step by changing one string, which is how a team can start with a large model for planning and drop to a smaller one for extraction once the pipeline works. Third, reading the credit dashboard so a team could see where its tokens were going before it ran dry mid-demo.

The teams that used this split were easy to spot at demo time. Their agents felt fast because the branching was fast, and their multimodal features worked because image and video calls went to models built for them.
Twelve people scored the demos. The panel mixed model labs, dev tools, founders and investors, which made the debate at the table more useful than a single-lens jury would have been.

Judge | Role |
|---|---|
Ashley Khoo | Member of Technical Staff, Cognition |
Lucas Gonzalez Pagliere | Research PM, Anthropic |
Qasim Wani | Member of Technical Staff, xAI |
Sasha Zhang | Co-founder and CEO, Scout |
Gabriel Jarrosson | General Partner, Lobster Capital |
Xinchi Qi | Co-founder and President, Revamp |
Hendrik Krack | DevEx Engineer, CodeRabbit; Co-founder, N-aible |
Claire Xie | Founder, Women in AI Club |
Ken Morimoto | Investor, Leading Edge VC |
Julie Chen | Head of Marketing, Photon |
Sourabh Mane | Design Engineer, CodeRabbit |
Roan Weigert | Developer Relations Lead, GMI Cloud |


55 teams submitted a working project. Each judge scored every submission, the highest-scoring teams demoed live in front of the room, and the panel picked the winners from those. The judging criteria put weight on novelty, which Allie Laabs, who runs DevRel at TypeSafe, said was the thing she cared about most: an old idea that was infeasible before Jev and possible now.



The winner was Jevolution, a three-person team that built an evolution simulator in three hours. Rabbits and wolves with evolving genomes, every animal's next move decided by Jev, with a large model only setting up the experiment. Left to run, the rabbits developed a form of signalling between themselves and the wolves started hunting in packs, with zero code written for either behavior. The pitch was synthetic behavioral data for endangered species that researchers cannot experiment on; a 2022 study they cited estimated that 56% of 7,699 data-deficient species are threatened. What won it for the judges: the Jev approach unlocked a middle path between a rules-based simulation and an LLM agent per animal, and that path was something net new.
A few other projects worth knowing about, picked because they show the decide-then-generate split working in different domains:
FirstMinute: a paramedic describes a patient in plain speech, Jev scores the RACE stroke scale item by item and asks back only about the items it is unsure of, Browserbase checks live hospital status, and the SBAR handoff for the stroke team gets written by a model on GMI Cloud. Their own summary of the pattern: Jev scores, code decides, the medic confirms.
KYC Sentinel: adverse-media screening for compliance analysts. Jev judges every client and article pair (same person? the perpetrator? how severe?), a unit-tested router sorts clients into clear, review or flagged, and the LLM only writes the summary. Any uncertainty fails closed to review.
JevBag and Flinch: two separate teams had the same idea, a hook that asks Jev a handful of risk questions about every shell command or tool call before a coding agent runs it. JevBag replayed a day of real Claude Code commands and caught an rm -rf on a photo catalog, with each check taking about a tenth of a second.
LiftLine: a team whose badge failed in the elevator on the way to the rooftop built a phone intake agent for elevator service calls. Entrapment and injury route straight to a human, routine maintenance flows through Jev decisions.
Brailly: a Chrome extension and Braille display simulator that reads a webpage in the order that matters for the task, with Jev ranking an accessible entrance above the gift shop promotion.
Jevfusion and Jiffusion: two teams tried image diffusion in pixel space using nothing but Jev decisions, starting from noise and hoping for a circle. One of them met at the breakfast table that morning.
Beethoven deserves its own mention: paintings arranged on a musical staff, Jev conducting every bar, and a music model on GMI Cloud playing it live. Zero pre-recorded audio.
The teams that shipped scoped down in the first twenty minutes. One agent, one job, one clear input and one clear output. The teams that struggled tried to build a platform in three hours.
Jev changed how people structured their agents. In a normal LLM agent the model both decides and writes, and every decision costs a full generation. With Jev in the loop, teams pulled the branching out into typed calls and kept the generative model for the steps that produce something a human reads or watches. The agents that did this had visibly fewer LLM calls, and their demos moved.
Model choice was visible on stage. A few teams put the biggest model on every generation step and ran out of patience waiting for their own agent. The ones who matched model size to the step, a big reasoning model for the plan and a small one for extraction, had the snappiest demos.
Multimodal was the differentiator. Anyone who fed a screenshot, a photo or a video frame into the loop, let Jev score or classify it, then generated from the result, had a demo that felt new, because the audience could see the input and the output side by side.
Credits removed a real blocker. Several people told me they would have skipped the video and voice steps entirely if the calls had come out of their own pocket. Giving teams inference credits changed what they were willing to try in three hours, and that showed on stage.
The GMI Cloud MaaS inference engine has the same text, vision, image, video, and voice models the teams called on Saturday, behind one OpenAI-compatible endpoint, and the model quickstarts walk through each one. New accounts get free credits to try the same split on their own agent.
Thanks to TypeSafe AI and The AI Collective for co-hosting, to CodeRabbit for co-organizing, to HackerSquad for running submissions and judging, and to the eleven judges who gave up a Saturday.
Roan Weigert
DevRel Lead @ GMI Cloud
GMI Cloud helps you architect, deploy, optimize, and scale your AI strategies
