osENV.io open to work
Contactme

rendered in blender · osenv 0.3

I build judges for machines that write code.

osenv puts a crew of cheap AI coding agents to work on your machine, and a judge on every move they make. Before any of them writes a file or runs a command, the judge scores that action against the job, your rules, and every mistake the crew has ever made. Known mistakes are caught in the act. Nothing counts as done without proof.

One Go binary, no dependencies. Built solo, and put through 11 release candidates on Linux and Windows before it shipped.

proof, not promises

Caught in the act

board · task ledger · Linux test run2026-09-25 UTC
01:14:01
muse
writes ledger/ledger.py, its first draft of a CSV totals tool
jev
KNOWN MISTAKE IN THIS ACTION
L85 0.95 money summed and rounded as binary floats
L6  0.92 a port number parsed with a bare int()
osenv
sent back once: fix it in this write
01:14:15
muse
rewrites it with Decimal and a strict digit check before int()
jev
ALLOW

A real run. The worker's first draft repeated two mistakes that earlier runs had turned into lessons. osenv sent the write back, and 14 seconds later both were fixed, before any reviewer saw the code.

by the numbers

85lessons carried from project to project, learned in real runs on Linux and Windows
3 / 3send-backs were real bugs, replaying 80 real file writes. Zero false alarms.
0.1 sper judge call. 214 real actions re-judged in 18 seconds.
75tests on every build, on Linux under the race detector and on Windows

how it works

The crew works. The judge watches. The lessons stay.

  1. You, or your own Claude acting as the desk, write the job and the gates it must pass.
  2. A cheap worker model does the work, and every file write and command goes past the judge first.
  3. Reviewers step in for one run when the job asks: one model for code, one for how things look.
  4. Proof or it didn't happen: real test transcripts, planted-bug checks, live server checks, screenshots actually opened.
  5. Every mistake becomes a lesson, checked against every action from then on.

let's talk

Hiring for AI research, agents or dev tooling? I want to hear it.

Roles, collaborations, or what you'd build with osenv: this goes straight to me.