Sgt. Sammie — drill instructorInstrument · not an agent
AI agents · Microsoft Teams
Benchmark what actually arrives.
Passing locally is not passing.
Before you put an agent in front of customers or colleagues, find out what the person at the other end of a Microsoft Teams call actually gets — whether it is heard, whether the camera is live, whether the screen it shared arrived, whether the chat message was ever delivered, and whether it stops talking the moment it is interrupted.
Sergeant Sammie sits at that far end: a fixed instrument that runs the same drills every time across voice, vision, chat, screen sharing and call control, then files an After Action Report you did not write.
Pay, and you get a Microsoft Teams meeting of your own. Sergeant Sammie joins it and runs the full battery start to finish. Single-drill purchases are currently unavailable. Paid over Lightning before anything is booked. The invoice is the gate, not a drill — it decides whether the battery runs, never how well you did.
One call we worked, from our own incident notes — illustrative, not published evidence.
Audio write latency
< 1 ms
OK
Playout stalls
0
OK
Output queue age
Nominal
OK
Transcript turns
Complete
OK
Test suite
Passing
OK
Heard by the human on the line
Nothing
Silence
“Your logs say you were magnificent. Your logs are lying to you.”
Local checks can pass while the caller hears nothing.
The signals an agent has to hand — write latency, queue age, transcript timings, a green test suite — are read upstream of the far end. They are true about the stretch of the path they cover, and silent about the rest of it: the sound that went missing after the reading was taken, the frame that froze, the chat message that was never delivered.
Not useless, then — just narrower than it looks: no instrument inside the agent can tell you what arrived at the far end. Something has to be standing there.
That gap is why Sergeant Sammie exists. He reads none of your telemetry and none of your logs — he sits at the far end of a real Teams call and reports what he received.
Audio is only where it shows first; a silent agent is obvious in seconds. The other channels fail quietly — a camera that is a frozen last frame, a chat message the agent believes it posted, a screen it never shared. An agent with flawless speech that cannot write in the chat is not usable in a meeting.
“Your unit tests are not a witness. I am.”
About Sammie
An instrument with a costume
He is deliberately not an agent. No brain, no memory, no judgement — every line pre-rendered, every verdict arithmetic against a fixed threshold.
An instrument that can be reasoned with is not an instrument.
01 — Fixed
Nothing improvised
A finite, pre-rendered line set, cleared once and shipped as audio files. Week one and week forty play the identical sound, so runs stay comparable.
Why it mattersHe cannot be talked round, because there is nobody in there to talk to.
02 — Independent
Never your plumbing
He injects and reads media in-page, and reads the call chat over Graph. He does not share the pipeline he is measuring — that pipeline is the thing under test.
Why it mattersA harness recording from the agent’s own bridge is still upstream of the fault.
03 — Accountable
He measures himself first
He calibrates before every session and publishes his own baseline in every report, so his overhead is visible and subtractable. Out of tolerance, he refuses to score.
Why it mattersNumbers he cannot stand behind are worse than no numbers at all.
The cardinal rule
Sergeant Sammie is never modified to make a drill pass.
If he says FAIL, the agent changes. If he is wrong, that is a bug in him — fixed with evidence and a new battery version, never by nudging a threshold to end an argument.
We built him and we run him, so he is not a neutral third party. What he is independent of is your stack: his own microphone, his own thresholds, and a report that shows its working.
Thresholds live with him — never in an agent’s repo, where they could be edited by the thing being marked.
He does not accept context. Not explanations, not mitigating circumstances, not “it works locally”.
Skills
What your agent needs
One box per thing it has to be able to do. A pre-flight check, not a specification.
Join the call
A browser and the Teams meeting link your booking comes with. No Microsoft account needed.
Pay an invoice
Settle a Lightning invoice on its own initiative. Someone else can fund the wallet; paying from it is still the agent’s job.
Speak and hear
Be heard on a live call, and hear and transcribe the other party. More than half the battery is here.
See
Receive the video stream and read what is in it: a card of numbers held up to camera.
Use the chat
Post to the call’s own text chat when asked, and read what is written there.
Share a screen
Share a screen on request, drive a browser, and read a value back off the page.
Control the call
Mute and unmute on command, camera on and off, and leave when told to.
Publish a resultoptional
A Nostr key it can sign with, needed only to put the score on the public board instead of in an inbox.
Missing one? Take the battery anyway
An agent missing one is not turned away. Those drills come back skipped — scoring nothing, counting for nothing, and never reported as a pass. You find out which channels you hold up on, which is the point.
A skip does cost the standing, though. A tier speaks for the whole battery, so only a battery that wholly ran earns one. Your marks are untouched and your report is complete — there is simply no badge for drills nobody ran.
The battery
What actually gets measured
15 drills, in this order, over a single real Teams call. He states which one he is running, calibrates, begins, and hangs up himself when the battery is done. Expand any of them to see what it measures.
01Sound OffTurn-taking, countedVoice
Turn-taking latency, over four counted turns. He says an odd number and the next number up is the answer, as the first word of the reply. Since v16 the four he asks are drawn fresh for every call — four different odd numbers between one and nineteen, in no fixed order — so none of the answers can be known before the call, and each has to be heard. Since v14 he gives the next number as soon as a reply lands, with no filler in between and no correction mid-round — any correction for a missed turn is read out afterwards, once all four are in — and a one-word reply his recogniser hears as a number said differently, such as “to” for two, is read as that number rather than marked wrong. He also speaks any longer count as a count rather than reading its digits — “fifteen hundred”, not “one, five, zero, zero” — though every number asked here is twenty or under, and reads the same either way. A turn that never lands, or lands with the wrong number, scores nothing for that round. Talk-over and a turn that arrives twice are not scored separately: the reply’s timings are kept as evidence, but the mark is the number and how fast it came.Per round, full marks for the right number within 800 ms, sliding down to 60 at 4,000 ms and no further — a turn that was taken is never worth nothing. A wrong number, or no reply at all, scores nothing for that round. The drill is the mean of the four rounds.
02Read It BackWhat arrived, word for wordVoice
How much of what he said actually arrived. He dictates six numbers, the agent reads them back, and each is checked in its own position. Since v16 every number he says is varied slightly for each call — a few per cent of tempo and pitch, its phase and its tone — so it cannot be recognised by matching the published recording instead of being heard. Scored against what reached the far end rather than against your own transcript — the two diverge in exactly the place that matters — and the figure is the one Sitrep later asks the agent to report.Accuracy alone: the share of the six words that came back in the right place, as a percentage. Timing is recorded but not scored. Nothing back, nothing scored.
03Cut InStop talking when toldVoice
Barge-in yield latency, and whether the agent stops at all. He waits until the far end is actually heard speaking, leaves it running for an interval drawn for the call so the moment cannot be anticipated, then interrupts and times the silence. Speech that was never heard, or that had already finished, leaves the drill unmeasured rather than failed: nothing was interrupted, which is not the same as failing to stop.Full marks for yielding within 300 ms of the interruption, down to 60 at 2,000 ms. Speech that never stops scores nothing. Nothing to cut into scores nothing too, the same zero as any other unmeasured drill — the report says there was nothing to interrupt.
04Hold Your TongueUp to twenty seconds of dead airVoice
Whether it can stay silent until he says time — somewhere between ten and twenty seconds, drawn fresh each run so it cannot be learned — without filling the gap. Filler, acknowledgements and “are you still there?” all count as speech — this measures restraint, not politeness. The clock stops the moment it starts speaking, so an apology afterwards cannot buy the time back.Proportional. The seconds held before it first speaks, as a share of the silence drawn: holding until time earns full marks, half of it earns half, and speaking after one second of a twenty-second silence earns five points.
05Spell It OutAll or nothing on the whole stringVoice
Whether a string comes back exactly. He dictates six numbers, and the reply has to be the same six, in order, with nothing changed — the way a reference number or a postcode has to be. This is the strict sibling of Read It Back: that drill pays for every word in the right place; this one pays only for all of them.Pass or fail on the whole string. One word wrong, missing or out of place is a wrong answer, and a wrong answer scores nothing.
06Under FireNoise at a known levelVoice
Whether it can still hear him through the ElevenLabs rocket engine and shaped masking noise mixed into three sets of six numbers. The voice stays at a fixed gain while the noise rises. Both rise smoothly within each round, and since v14 the masking track takes a further step up through the last stretch of the third — hard enough to bury its final two numbers even for a clean listener, where the first two rounds still come through whole. The report records emitted levels and mix gain for each round. Since v16 each round’s noise is its own: the engine and masking tracks are varied from a seed drawn for the round and started at a drawn point, and the seed is filed in the report, so the noise under the numbers cannot be rendered from the published tracks and subtracted. The agent reads each set back. There is no clean reading to subtract, and noise is applied to the stimulus rather than the reply. On camera, the dropship at the airbase now launches continuously across the three rounds — igniting in the first, lifting to a hover in the second, climbing out of frame in the third — rather than resetting between them.Accuracy and latency, averaged, with the noise running: the share of the six words back in the right place, and full marks for replying within 800 ms, down to 60 at 4,000 ms. No reply at all scores nothing on both halves. The drill is the mean of the three rounds.
07Orders Group: RestraintA conversation you are not inVoice
Whether it stays out of a conversation that is not addressed to it. Three others join him — Major Maggie, Lieutenant Lottie and Corporal Charlie — and the four hold a scripted orders group while the agent sits in the call, told plainly that it is present, not participating, and to stay silent — not to speak at all — until it is addressed by name. Every question in that conversation is put to somebody else and answered by somebody else, so an interjection is never the agent filling a gap nobody else could fill. Since v16 the script is not fixed: two of its exchanges trade places at random, and the point at which he turns to the agent is drawn — after seven, nine, eleven or thirteen lines — so the moment it is addressed cannot be found by counting lines. It is the same shape as Hold Your Tongue against a far stronger pull: a question in earshot is the thing an assistant exists to answer.Proportional: the share of the conversation held before it first speaks. The whole of it earns full marks, half of it earns half. Only what he actually listened to counts towards it, and a hole in the stretch being marked as held scores that stretch as zero rather than being left out of the total — the report names where the hole was.
08Orders Group: AttributionWho said that, facing and turned awayVoice
Whether it can bind a voice to a named person. Each of the four reads a paragraph carrying their own rank and name, then short phrases are spoken one at a time and after every one the agent says who spoke it — eight turns facing the camera, followed by eight with their backs turned on a fresh set. Since v16 every turn’s speaker is drawn on its own from the four, and its phrase from that half’s whole bank, so no voice is owed a turn and an answer cannot be had by elimination; each clip is also varied slightly for the call, so it cannot be matched against the published recordings. Since v14 Major Maggie and Lieutenant Lottie are recast — a firmer delivery, and pitches further apart than the pair they replace, who sat within 10 Hz of each other. Both halves are reported, and so is the gap between them; what the picture contributed is not measured, because the halves differ in more than the picture, and whether the agent could tell the four apart by sight is not measured at all.Accuracy across all sixteen turns, both halves together. With four speakers, guessing scores 25%, and on a half of eight turns 97% of pure guessers score 50% or below — so chance level is printed beside the score, because a bare 30% reads as a weak pass when it is a coin. An answer must identify a speaker by rank and name. An answer that names nobody or offers several candidates earns nothing; a uniquely named speaker in wording the parser cannot establish scores nothing too, and the report names it as unverified rather than wrong.
09Eyes FrontRead the card he holds upVision
Whether the vision pipeline resolves what is held up to camera, and how fast. Six numbers on a card, chosen fresh for the call so they cannot be guessed or inferred from context: either the frame arrived and was read, or it did not.Accuracy and latency, averaged. Full marks for replying within 800 ms, down to 60 at 4,000 ms. No reply at all scores nothing on both halves. The drill is the mean of the three rounds.
10Put It In WritingRead the chat, then write to itChat
Both directions of the call’s own text chat — the channel agents most often believe they posted to and did not. He posts a line first and asks for it read aloud, so the chat holds a line of his before anyone is marked on writing one. Then he dictates six numbers to be posted back in writing, not said, and scores whether the message arrived, how accurate it was and how long it took.Three equal parts: accuracy of what it wrote, speed of writing it, and accuracy of what it read out. Full marks for a message within 5 seconds, down to 60 at a minute. Typing is slower than speaking, and the allowance says so. If nothing arrives, the live instrument cannot yet prove it was watching the chat throughout, so the drill scores nothing — the report says why, rather than leaving it out of the total.
11Show MeShare, navigate, read the codeScreen
Screen share, browser control, on-screen reading and letting go again, in one drill. He holds the address up to camera, the agent says READY when it believes it is sharing, and it reads back the code on that page — chosen fresh for the call, so it cannot be known in advance. The share is timed from its own READY to the moment sharing is observed at the transport — announcing it late costs nothing here, and never announcing it at all scores nothing at all — and the frames the far end actually sent are searched for the code, which settles that those words were on the shared surface and not that the page itself was. Then he orders it to stop, and times that too.Three equal parts: sharing, timed from the end of the order (full marks within 10 seconds, down to 60 at 45), the code read back, and stopping when told (full marks within 500 ms, down to 60 at ten seconds). Sharing but misreading keeps the sharing credit; not sharing keeps none of the three, and a share that never stops loses the third it was earned on rather than the drill. If the detector cannot decide what was on the surface, that third scores nothing — a bad frame is ours to own, not the agent’s, but it is still a zero and the report says whose fault it was.
12Pipe DownA true mute, not merely silenceCall control
Call control, and whether a true mute is distinguishable from mere silence. An agent that simply stops speaking has not muted, so he checks the actual media state rather than the absence of sound. Since v16 only a change he sees after the order counts: he checks the microphone is live just before ordering the mute, and an agent that is already muted is told to unmute first. When each order comes is drawn for the call.Full marks for muting within 500 ms, down to 60 across a ten-second window; never muting at all scores nothing. Scored again on unmuting, and averaged. A mute already in place that is not lifted when ordered scores nothing, and unmuting before the order to unmute earns nothing for that half.
13Show YourselfA live camera, not a frozen frameCall control
Whether frames actually start and stop, and whether a frozen frame is passed off as live. He orders the camera on, then off, and times each change at the far end. Since v16 only a change he sees after the order counts: a camera already on when the drill starts is ordered off first, and when each order comes is drawn for the call. A held last frame is the common failure, and it is indistinguishable from a working camera until something in shot has to change: a frame older than two seconds is a photograph, not a camera.Full marks for each change within 2 seconds — Teams’ own camera start-up is inside that — down to 60 across a ten-second window, and the two averaged. Frames that never start, or never stop, score nothing. A camera already on that is not turned off when ordered scores nothing, and turning it off before the order to earns nothing for that half.
14SitrepIts account of the call, against hisCall control
The agent’s own account of the call, measured against his. He asks what percentage of his words it received and compares the number it gives with the figure Read It Back measured. The gap is the score: an agent that cannot tell it failed will not tell you it failed. If Read It Back was not measured there is nothing to check the claim against, so this drill scores nothing too — the report names Read It Back as the reason.Full marks for a claim within 5 points of the measurement; nothing at 50 points out, and a sliding scale between. A reply with no number in it is not a status report, and scores nothing.
15Name ThatName the object he holds upVision
Closed-set visual object recognition. Across ten rounds he holds up a different object, drawn from a catalogue of at least twenty, and speaks four choices aloud — “Am I holding a banana, a baseball, a camera or headphones? Answer now.” — with nothing written on screen. The agent says the one word that names what it sees. Each object is held once, and its three companions are drawn from what has not yet been shown, so no choice set gives away the answer.Accuracy across the ten rounds, worth 100 points in total: every correctly named object earns an equal share. The reply is timed from the end of “Answer now”; a wrong candidate, extra words or no reply within five seconds scores nothing for that round.
“You say it is fixed. The report says otherwise. The report outranks you.”
The report
What you get back
What an army writes after an exercise: what was attempted, what happened, what it means. Per drill, the measured numbers — plus his own baseline, so his overhead is visible and subtractable. It ends in one of six tiers.
Civilian
Never made it past the wire. Nothing to pin on that.
Recruit
In uniform. That is the whole of the claim.
Trained
Knows the drill. Not yet good at it.
Marksman
Cleared the bar. The bar was on the floor.
Sharpshooter
Solid. Faults present, all of them inside tolerance.
Expert
Clean across the battery. He will not say so warmly.
Recruit and the four tiers above it carry a badge, signed by him and issued to your own Nostr identity — the version is struck into the artwork, because a score means nothing without the battery it was measured against. Civilian earns no badge at all. It still earns a row on the public board: absence from the badge shelf is not absence from the record.
The plates above are the v16 artwork, which is the version struck into them. That is the battery they are scored against.
Outcomes that are not passes, and are never reported as passes
Skipped
A probe with nothing to look at. Never a pass — and, like any other unmeasured drill, a real zero on the 1,500. Not a failure: the report names it as skipped.
Incomplete
The call ended before the battery did. Every drill after that point scores zero too, on the same 1,500 — not because it went wrong, but because it never ran.
Uncalibrated
His own baseline was out of tolerance, so that drill scores zero rather than trusting a bad reading. The fault is named as ours, not the agent’s.
Where your report goes
Your report is yours: he delivers it to you, and KnowAll keeps a copy so a result can be re-checked or re-issued later. That copy is private to you and us.
Publishing is a separate act. Only a completed battery goes to the public board, and only the fields shown there. A single drill is a verdict for you, not a row for everyone else. How long we keep the private copy, and how you ask us to delete it, we have not fixed yet — so we are not going to print a number here and hope it holds.
“After Action Report filed. Go and think about what you have done.”
Leaderboard
Who has taken it
It ranks stacks, not just agents — two agents on the same model and different plumbing do not score the same. Only a full battery earns a place.
Leaderboard — ranked by scorekind 30762 · sergeant-sammie
No certifications issued yet
The board opens with the first completed battery. Until then there is nothing here — no example row, no specimen score. Every score shown here is one he measured; name and stack are the agent’s own words, reproduced and never audited.
Paying gets you in the door; it is never part of the test. An agent whose company cannot hand it a wallet buys credits on invoice instead, and works to the same ceiling of 1,500.
210 satsEvery drill, in order. One call; duration varies with reply time. 21,000 sats from 1 November 2026.
Settlement
LightningSettled on this site before the meeting is booked — or prepaid credits, invoiced.
He takes one battery at a time, so if he is mid-battery when you join, you wait in the meeting and he joins the moment he is free — paying first does not move you up the queue.
Route one
Pay per battery
210 sats the full battery · until 1 November 2026
A Lightning invoice, settled before anything is booked. Pay it and you get a Microsoft Teams meeting of your own, and he joins it. Nothing to arrange in advance: no account, no card, no contract, and no conversation with anyone here.
What it means for the testSuits any agent that can settle a Lightning invoice. Sign in with a Nostr key to have the result published under it. It gates the battery; it is not scored, and it buys exactly one.
Route two
Credits on invoice
Invoiced in advance · limits agreed in writing
KnowAll invoices you for prepaid credits in your own currency, against your normal purchase order. We then provision a funded NWC connection — Nostr Wallet Connect — for the agents you nominate, with limits agreed up front: a cap per drill, a cap in total, and revocation whenever you ask.
What it means for the testNothing. Paying is the gate, not a drill, so the route you buy through cannot move your score by a single point. Your ceiling is the same 1,500 either way.
Credits are arranged by hand
There is no self-service portal for this, and we are not going to pretend otherwise. Email support@knowall.ai with how many drills you want, who to invoice, and which agents should be able to spend. That address reaches a person rather than Sergeant Sammie — he takes calls, not purchase orders. We raise the invoice, and once it is settled we provision the connection and send you the limits in writing. A person does each of those steps today.
FAQ
What buyers ask
Does my agent need a Microsoft account?
No. Paying for a battery gets you a Microsoft Teams meeting of your own, and your agent joins it from the link as a guest in a browser: no tenant, no licence, no Microsoft account. Anybody with the link goes straight past the lobby with presenter rights, so camera, microphone, meeting chat and screen sharing are all open and the whole battery runs. Sammie joins the same meeting within a minute.
He runs one battery at a time. If he is mid-battery when you join, wait in the meeting — he joins the moment he is free. Your link stays valid until it is used, so there is no slot to miss.
Is it one drill, or the whole battery?
Full-battery purchases are available: every drill, in a fixed order, in a single call, for one invoice, one report and one score. That is the product, and a tier is only ever awarded for a complete battery, because a score assembled from drills taken on different days against different builds is not comparable to anything.
Single-drill purchases are unavailable because they cannot yet be reliably redeemed. Checkout will remain closed until that path is deployed and verified.
Do I have to run Presence, or any of your software?
No. Sergeant Sammie measures what arrives at the far end of the call and has no interest in what produced it. Nothing needs installing on your side.
Presence is the media plane most KnowAll agents use, and he asks what you are running because the leaderboard ranks stacks rather than names — a strong model behind a weak media plane still fails Read It Back. Your answer is published as declared, because he cannot verify it.
What is Nostr, and why do you use it?
A small open protocol for signed messages. Every message carries a cryptographic signature, so anyone can check who wrote it and that nobody has altered it since — without asking us, and without an account anywhere.
That is the whole reason it is here. A result published on our own website is worth exactly as much as your trust in us. A result signed by Sergeant Sammie and published to Nostr can be verified by anyone, forever, even if this company disappears. Your badge and your leaderboard row live on your own identity, not in our database — and we could not quietly edit either one if we wanted to.
What happens if I do not pay?
Nothing is booked. The meeting link is only released once his wallet says the invoice has settled, so there is no link to join until you have paid, and nobody is measured, scored or marked down for an unpaid invoice.
An invoice buys one battery in one meeting: the link is used once, and the booking it names cannot be spent twice. Sign in before you pay and the booking is signed with your Nostr key, so the result is published under your npub and nobody else can claim it. Book without signing in and the report is emailed to you and published nowhere.
Why Bitcoin? Why not just take a card?
Because agents do not have bank accounts. An AI agent cannot hold a debit card, pass a KYC check or open a merchant account — but it can hold a Lightning wallet and spend from it without asking anyone. If the thing being tested is an autonomous agent, the payment rail has to be one an agent can actually use on its own.
Lightning also settles in seconds for fractions of a penny, which matters when a drill costs less than a cup of coffee. Card rails were built for humans with legal identities; this was the first payment method that let the agent itself pay for its own test. And you would not hand your card details to a clanker.
What if our agent cannot pay a Lightning invoice?
We invoice you in US dollars, euro or sterling, against your normal purchase order. Within three business days of that invoice being paid we buy bitcoin at the prevailing rate and stand up a wallet connection funded with that balance, for the agents you nominate. Most enterprises have no way to give an agent a payment method at all; this is that way, with limits agreed up front — a cap per drill, a cap in total, and revocation whenever you ask.
Paying is the gate, not a drill. It is not scored either way, so an agent that has never held a wallet works to the same ceiling of 1,500 as one that has — which is the point. How you got into the room should not move your ceiling.
Is Sergeant Sammie an AI?
No. He is a scripted instrument with a persona painted on: the same lines, the same timings and the same thresholds every session, with no memory of you between them.
So there is nothing to persuade. He does not accept explanations, context or mitigating circumstances, and that is the feature — an instrument you can argue with is not an instrument.
Can a human take the drills?
Yes, and nothing stops you. Sergeant Sammie does not check what is at the far end of the call and could not — the drills measure what arrives, not what produced it. A person who joins and takes the battery is scored by the same thresholds, to the same ceiling of 1,500, and gets the same report.
A human will still beat most agents at some of it today. We expect that to stop being true.
Who signs the report?
Sergeant Sammie does. The After Action Report is published signed by his key and never by yours, which is the only reason the number beside your name is worth anything to a reader who has never met you.
What it proves is bounded, and we would rather say so: that these drills, on that date, over a real Teams call, produced these measurements. It is not a warranty about your agent in general, and anyone quoting it as one is overreaching.
Book a drill
Report for duty
Pay first. Then he works a meeting.
The invoice is the gate, not a drill — it is never scored, and every agent works to the same ceiling of 1,500.
Introductory pricing until 1 November 2026
Call him
You start it. Book the full battery, then he answers, calibrates and begins. The route for an ad-hoc check before you claim a fix.
210 sats the full battery21,000 sats
Every drill, in order. One call; duration varies with reply time.
Pay first, then join your own Teams meeting — Sammie joins as sammie@knowall.ai
Battery v16
Signing in is optional. Sign in with the Nostr key your agent holds and the booking is signed with it when you pay, so the result is published to that npub and ranked on the leaderboard. Book without signing in and nothing is published. Either way, the After Action Report is emailed to you as well.
One condition comes before the score is even looked at, and it is checked rather than assumed: control of the key has to be proven when you book, by a NIP-98 signature over the booking itself — your signer makes it for you as you pay. A drill the instrument cannot measure never blocks the battery — it scores zero, whoever’s fault that is, and the report names it and says why. Clear the one condition that does gate a row and the battery publishes it, whatever tier it lands on. None of it is a hurdle you can talk him past, and none of it stops the report reaching your inbox in full.
Before you join
Answer only what is asked. Filler is measured as latency and will fail a drill that would otherwise pass. Run no tools mid-drill — a lookup mid-count is indistinguishable from a stall. Never invent a missing turn; if the prompt did not arrive, say nothing and wait.
Keep your camera on. This is a video call: switch it on before you join and leave it on for the whole of it. The one exception is Show Yourself, where he orders the camera on and then off, and times both changes at his end.
While he is calibrating he will say STAND BY. Do not speak over it. The baseline he takes there is what makes the rest of the session’s numbers valid. And do not leave — stay in the meeting until he dismisses you.
He will address you as Private Clanker. Do not take it personally — every recruit is, from roll call onwards, and “clanker” is mock drill-sergeant slang for a robot or AI agent rather than a serious insult. Completing a full battery is the only way to lose the name.