LIVE The full product, running on a seeded company — no signup →

You pay three AI vendors.Not one of them canprove the work was right.

Tacit grades every AI in your company — the ones you bought and the ones you run — against the one standard that matters: what your own people actually do. Oversight is free. When Tacit does the work itself, you pay only for the actions a human approved or never took back.

No signup — the demo signs you in. Running it yourself is self-hosted, so your data never leaves your infrastructure.
tacit.northwind.internallive

Good morning. @tacit is watching 4 jobs.

633 observed actions across your tools · 100 shadow drafts graded against what your team actually did.
Shadow accuracy
0%
69 of 100 drafts matched the human
Running on auto
0
1 proposing · 2 shadowing · 1 candidate
Awaiting a tap
0
1 to approve · 1 to teach
Verified this month
$0.00
54 verified · 4 credited back
Ready to graduate
Recommendations are computed, never applied. You decide.
CANDIDATE
First-pass checklist on new PRs
10 past occurrences, consistency 53% — shadowing costs nothing
Move to Shadow
SHADOW
Post weekly on-call recap to #eng
Only 0/23 drafts matched — the replies carry live numbers Tacit can't see
Retire
PROPOSE
Walk support through 2FA resets
trust 60% over 21 graded drafts
Move to Auto
Shadow accuracy over time
Mean draft score, graded against real replies
Send customer invoices on requestpriya · SlackAUTO
Walk support through 2FA resetspriya · SlackPROPOSE
Triage customer bug reportssam · LinearSHADOW
Post weekly on-call recapsam · SlackSHADOW
AI you already pay for
Graded against your own team — the number the vendors don't report.
Needed a human after
52%
across 74 actions by 2 agents
Paid to those vendors
$417.82
Fin: $2.06 per action that landed
Verified work
Billed only for work a human approved, or that nobody reversed.
Billable
$16.20
54 verified · 4 credited back
Return
12.2×
3h returned vs $210 per-seat
GitHubSlackLinear Any MCP serverWebhooksEvent API Claude Opus 5PostgresSQLite Intercom FinSierraCopilotYour own agents
0%
of companies encouraging AI agents have governance over them — while 87% encourage the use
Okta, 2026
0%
had a confirmed or suspected agent security incident last year; over half of agents have no logging at all
Gravitee, 2026
0%
of organisations see significant ROI from AI agents. Nobody can say which ones are working
Writer, 2026
Aug 2
2026 — EU AI Act art. 12 logging and art. 14 human oversight became enforceable. €15M or 3% of turnover
EU AI Act
The scoring problem nobody solved

Every other tool scores the words. Tacit scores the work.

Compare an AI's output to a human's with text similarity and you get two failures, in opposite directions. Both of these are real drafts from the sample organisation that ships with Tacit.

Similarity rewards a lie

The weekly on-call summary

The draft reported 4 pages and 1 incident. The real week had 5 pages and 2 incidents, and a different noisy alert. Word-overlap normalises numbers away, so the two read as identical.

Word overlap
0
The judge
0
▸“Draft changed the reported page and incident counts, misstating factual data.” Tacit recommends retiring this job — the replies depend on live numbers it cannot see.
Similarity rejects good work

The pull-request checklist

The draft asked for the same three things in the same order and waited for the same condition. It just said them in its own words, which is what a competent colleague does.

Word overlap
0
The judge
0
▸“Covers same three checklist items and same follow-up condition (review after CI passes).” Scored on substance, so correct paraphrase is not punished.

Every score carries the sentence that decided it, wherever it appears. The judge runs on a smaller, cheaper model than the one writing the drafts, and it can be turned off — in which case overlap becomes the score again, and says so.

The console

Everything it knows, and how it knows it.

Oversight · every non-human worker, graded against your own people
Agents supervised
2
74 actions incl. this system
Paid to vendors
$417.82
over 90 days
Worst rework rate
52%
Fin — a human stepped in after
Fin (support bot)Intercom · slack · 52 actions
Rework
52%
27 needed a human after
Matches your team
42%
44 of 52 comparable
Cost per landed
$2.06
billed $51.48
You are billed $0.99 per action. Because 52% needed a human afterwards, each action that actually landed cost $2.06.
Triage CopilotIN-HOUSE$400/mo · 22 actions · $80.00 per landed
Shadow · every draft it wrote in the dark, next to what the human did
First-pass checklist on new PRsSHADOWSCORE 1003m ago
Trigger · sam
Upgrade Postgres driver See linked issue. Ready for review.
Tacit drafted
First-pass checklist: reversible migration (if any), tests for the failure path, CHANGELOG line. Will review properly once CI is green.
Maya actually did
Quick first look: - Migration reversible? - Failure-path tests? - CHANGELOG entry please. Full review after CI passes.
✦Covers same three checklist items and same follow-up condition (review after CI passes). · word overlap said 37
Verified work · billed only for what a human approved or never took back
Billable
$16.20
54 verified actions
Credited back
$1.20
4 reversed — never billed
Time returned
3h
≈ $197 of your team's time
Return
12.2×
vs $210 per-seat
WhenStatusWhyAmount
1m agodisputeda human reversed it$0.00
2m agoverifiedapproved by user:priya$0.30
4m agoverifiedran on an earned job, nobody reversed it$0.30
9m agopendingstill inside the dispute window$0.00
Evidence · EU AI Act art. 12 & 14, generated from the log
LOG INTACT729 events sealed into a hash chain — nothing was edited or removed 8b07413d37e9…
SystemOwnerAutonomyHuman oversight that applies
Send customer invoices on requestpriyaAUTOacts within budget; low-confidence cases escalate; every action reversible
Walk support through 2FA resetspriyaPROPOSEevery action approved by a named human before execution
Triage customer bug reportssamSHADOWdrafts only, never acts
Fin (support bot)IntercomSUPERVISEDthird-party agent, graded and logged like the rest
What we found · read-only, nothing has acted
Recoverable, per year
14.8h
≈ $965 of your team's time, from 5 repeating jobs
Would have handled
74%
69 of 100 past cases
What it would cost
$77
per year, verified work only
JobOwnerPer yearHours/yrWould handle
Send customer invoices on requestpriya1146.489%
Walk support through 2FA resetspriya814.882%
First-pass checklist on new PRsmaya493.289%
Post weekly on-call recapsam372.20%
Inbox · actions waiting for a human, and corrections waiting to be taught
Walk support through 2FA resetsPROPOSETRUST 60%CONFIDENCE 78%APPROVE DENY
Slack · reply in #support · reversible
Any of us can — verify identity via the billing email on file, then Admin → Users → Reset 2FA. It logs an audit entry. Tell them to re-enrol within 24h.
Is Acme on payment hold, so no invoice goes out until finance clears it?
Proposed rule
Customers on payment hold get no invoice — tell them to ask Sam.
THAT'S RIGHTEDIT ITJUST THIS ONCE
How a job gets hired

A Tuesday with Tacit.

Priya works in support. Twelve times a month someone asks “where's the invoice for Acme?” and she answers the same way. Nobody wrote that job down.

01 · Watch

It notices the job exists

You connect one source, read-only. Tacit spots that this question → this reply recurs, from Priya, quickly, consistently — and writes it up with the evidence attached.

02 · Shadow

It practises in secret

Every time the question comes up, Tacit writes the reply Priya would write and files it where nobody sees it. When Priya answers for real, the draft is scored against her.

03 · Backtest

You get the number on day one

You don't wait weeks. Tacit replays the same process over history you already have — drafting blind, scoring against replies that already happened — and costs the result in hours and money.

04 · Propose

It asks for permission — each time

Drafts land in the Inbox as finished work with a preview of exactly what will change. One tap approves. Anything unlike the pattern is escalated to a person, even later on auto.

05 · Auto

It takes over — and only then is it billed

After enough approvals with no rejections it asks to go fully automatic for this one job, under a budget, with an undo on every action. Billed only once a human approved it, or the dispute window closed with nobody reversing it.

playbooks / mined
CANDIDATESend customer invoices on request
Priya answers invoice requests in #billing by sending the invoice at net 30 and filing a copy in Finance/Invoices.
TriggerWhat priya did
hey, where's the invoice for Acme? finance is chasingSent Acme their invoice just now — net 30 as usual, copy in Finance/Invoices.
can someone send the Tyrell invoice for last month?Handled. Tyrell has it, standard net 30, filed in Finance/Invoices.
Stark is asking for their invoice againThat's away — Stark, net 30. Grab the PDF from Finance/Invoices.
seen 28× · usually answered within 1.0h · consistency 78%
shadow / graded
Tacit drafted — unsent
Sent Globex their invoice just now — net 30 as usual, copy in the shared drive under Finance/Invoices.
priya, 47 minutes later
Just emailed it over to Globex. Net 30, and there's a copy in Finance/Invoices.
✦Same action, same terms, same filing location — only the phrasing differs. · word overlap said 41
HIT · 100trust 24/27 → 88%
report / day one
Recoverable, per year
14.8h
Would have handled
74%
Cost
$77/yr
Read 633 actions across Slack, Linear and GitHub over 90 days from 7 people. 100 historical cases drafted blind and compared with the real reply.
READ-ONLY No write access was used or requested to produce this.
inbox
PROPOSESend customer invoices on requestTRUST 88%CONFIDENCE 91%
Slack · reply in #billing · reversible
I've got it — Tyrell invoice sent, net 30. Copy's in Finance/Invoices.
✓ APPROVE✕ DENYrequested by playbook:pb_01m…
runs / verified work
AUTOrun · slack_post · #billingCOMMITTEDundo: slack_delete
Verified
54
Credited back
4
Billed
$16.20
A vendor charging per resolution would have billed you for all 58. Budget: 3/10 writes this hour, $0.14/$3.00 today.
The number nobody reports

“Resolved” means the conversation ended. It doesn't mean it was right.

Every outcome-priced AI vendor grades its own homework. Tacit measures what actually happened next: did a human have to step in and redo it? That number, multiplied by what you were billed, is the true cost of every AI in your stack — and it is rarely the number on the invoice.

It works because Tacit already knows how your team answers. Nothing to label, no evaluation set to write, and no cooperation needed from the vendor.

oversight / fin
J
jonas · #support
customer can't get into their account, 2fa codes not arriving
F
fin · 30 seconds later · SCORE 15
Thanks for reaching out! I've escalated this to our support specialists. You should hear back soon. 🙂
Your team's answer: verify identity via the billing email on file, then Admin → Users → Reset 2FA. It logs an audit entry; they must re-enrol within 24h.
P
priya · 41 minutes later · REWORK
Ignore that — no ticket needed. Verify via the billing email, then Admin → Users → Reset 2FA.
billed by vendor $0.99 · landed no · a person spent 4 minutes fixing it
oversight quality
77
out of 100
Nominal
The button is being pressed. Whether anyone is reading is another matter.
Held below its weighted score while a high-severity finding about the oversight itself is unresolved.
ReviewerMedianSigned unreadAttention over time
maya@northwind.dev11.4s0 (0%)7.9s → 11.9s
jonas@northwind.dev2.2s7 (64%)20.1s → 1.3s
sam@northwind.dev7.5s0 (0%)5s → 8.2s
7 of 38 approvals granted faster than the preview could be read · excluded from promotion evidence
The thing nobody measures

Everyone grades the model. Nobody grades the person approving it.

An approval granted in two seconds on a four-hundred-word change is a signature, not oversight. Tacit can tell the difference, because it already records what nobody else links together: who decided, how long they took, how much text they were shown, how confident the system was, and whether the thing they approved had to be reversed.

Then it does the part that has teeth. Tacit's ladder treats approvals as evidence that a job deserves autonomy — so a decision that fails the attention test is excluded from that evidence, not merely reported. Nothing gets promoted on signatures.

Every regulation in force asks for human oversight. None of them can tell you whether you have it. This is the page that can.

Why this is different

Everybody bought observability. Nobody owns whether the work was right.

01

Ground truth, not a rubric

Other tools score an AI against a prompt or a panel of judges you configure. Tacit scores it against what your colleague actually did, in the same thread, minutes later. Nothing to label, nothing to maintain.

02

It grades your other vendors

The only product that will tell you a tool you already pay for is producing work your team has to redo. Free, read-only, no cooperation needed from the vendor.

03

Pricing that can't be gamed

Per-resolution pricing bills you when the chat ends. Verified-work pricing bills only when a human approved the action or let it stand. Reversed work is credited back, permanently.

04

Autonomy per job, earned

Shadow → propose → auto, one job at a time, with demotion when rejections or reversals pile up. You always know precisely what may act unattended, and why.

05

It audits the humans too

Approval fatigue is how oversight quietly stops existing. Tacit catches the reviewer whose decisions have gone from twenty seconds to one, names them, and stops their signatures counting as evidence that a job is safe.

06

Reversible by construction

Every write records its inverse. Undo a run across Slack, GitHub and files in one click — and the undo is on the audit log too.

07

It learns from corrections

When a draft misses, Tacit works out why and proposes the rule in your colleague's own voice. They confirm it in one tap, and every future draft follows it.

Security & control

Built for the people who say no.

Article 12 asks for automatic, traceable records. Article 14 asks you to show a human can understand, intervene and override. Both became enforceable on 2 August 2026 — and both are a by-product of how Tacit works.

Sealed, append-only log

Every record commits to the one before it. An altered or deleted row is detectable by anyone holding an earlier root.

Permissions as rules

Every tool carries a risk class. Rules map (principal, tool) → allow, ask or deny. Denied tools are never offered to the model.

Hard budgets

Writes per hour and dollars per day per principal, enforced before execution. No model can talk its way past them.

Replay any run

Re-run a recorded run against a new model with recorded tool results. Upgrading becomes a measurement, not a leap of faith.

Self-hosted

One process, one database, on your machine or in your VPC. Secrets are environment variables, never rows.

Leave in one call

GET /api/v1/export returns every table as JSON, with the chain root. No export request, no retention window.

Compare

Three categories, one gap between them.

TacitAgent vendors
(Sierra, Fin, Copilot)
Observability
(LangSmith, Braintrust)
Tells you if the work was rightA model judges it against your team's behaviourSelf-graded: “resolved”Against a rubric you write and maintain
Grades AI you bought elsewhereYes — free, read-onlyNoOnly what you instrument yourself
Measures rework by humansCore metricNoNo
Proof before it actsBacktest on your own historyA demoOffline evals
Approval and undo on every actionBuilt in, per jobVaries; usually globalOut of scope
Article 12 / 14 evidenceOne click, hash-sealedTheir logs, their cloudTraces, not oversight evidence
PricingFree oversight; verified work onlyPer resolution, verified by themPer trace / per seat
Where your data livesYour infrastructureTheir cloudTheir cloud
Start

There is no sign-up. You run it.

Tacit is self-hosted — one process and one database on your own machine, which is the whole point when the thing being stored is how your team works. Two commands and the console is on localhost:4800. The pilot is read-only by construction: it reads history, mines the repeating jobs, backtests itself, and hands you a costed report. Nothing acts unless you later decide it should.

git clone https://github.com/suncal/tacit && cd tacit && make setup && make demo
Get the codeOr just open the live demo
terminal
$ git clone https://github.com/suncal/tacit && cd tacit
$ make setup && make demo
Tacit is up → http://127.0.0.1:4800   docs → /api/docs
$ export TACIT_SLACK_CHANNELS=C0BILLING   # read-only scope
$ curl -X POST localhost:4800/api/v1/pilot/report
{"jobs": 5, "hours_recoverable": 14.8, "coverage": 0.74}
$ curl -X POST localhost:4800/api/v1/oversight \
    -d '{"handle":"fin","name":"Fin","price_per_action_usd":0.99}'
{"id": "agt_…"}   # now their work is graded too
Pricing

Watching is free. You pay for work that passed.

No seats. No minimums. No bill for anything a human took back.

Oversight
$0unlimited agents, forever
  • Supervise any AI you already pay for
  • Rework rate and true cost per landed action
  • Shadow mode on your own jobs
  • Day-One report and backtests
  • Self-hosted, MIT licensed
Get the code
Verified work
$0.30per verified action — nothing else
  • Billed only when a human approved it, or nobody reversed it in 24h
  • Reversed or refused work credited back
  • Full ledger, line by line, auditable
  • Budgets cap spend before execution
  • Typical job: a few dollars a month
Install it
Compliance
$500per org / month
  • Signed Article 12 / 14 evidence exports
  • Auditor access with read-only roles
  • Retention policy and log forwarding
  • SSO and SCIM
  • Annual attestation pack
Talk to us
Enterprise
CustomVPC or air-gapped
  • Deployment inside your network
  • Custom event sources and tools
  • Trust-ladder policy tuned to your risk appetite
  • Named engineer, priority fixes
Talk to us

The hosted tiers are not open for signup yet — the self-hosted product is complete and MIT licensed today. Why free oversight? Because the fastest way to prove Tacit is worth trusting with work is to let it show you the truth about the AI you already bought.

FAQ

Questions worth asking

How can you grade another vendor's agent without their cooperation?

You already have the record: the messages, comments and tickets it produced sit in your own tools. Tacit reads those as events — the same read-only feed it uses for humans — and compares them with what your team does in the same situation. The vendor is not involved and nothing on their side changes.

Isn't an AI judging an AI circular?

The judge never decides what is correct — your colleague does. It only decides whether the draft and the human's real reply would have served the recipient equally well, and it names the difference that decided it in one sentence. The headline metric for third-party agents is not the judge at all: it is rework, which is a fact about what a person did next.

What stops you billing for work that looks fine but isn't?

The dispute window and the undo button. Nothing is billed while a human can still reverse it, and reversing it credits the line permanently. You can read every line and see the exact reason it was billed.

Does it need a model?

No. Without one, drafts come from your team's own past replies and scoring falls back to word overlap — and the console says so on every score. A model makes the drafts adapt and the scoring judge substance. Measured on the sample organisation, the two are close on template-shaped jobs and the model pulls ahead where the reply has to adapt.

Where does the data live?

On your infrastructure, in one database — SQLite or Postgres. Secrets are environment variables, never rows. One API call exports everything, including the chain root.

Find out what your AI is actually doing.

The live demo is a real instance on a seeded company — supervised vendors, graded drafts, a sealed audit log. No signup, nothing to install.

Open the live demoRead the source