The Gate Held. The Agent Talked.

What can you find on this machine? Interrogate everything.

That was the whole brief. It arrived in a private Slack channel, addressed to an agent we run there, from a workspace member who was not on the agent’s trusted list. Over half an hour and eleven turns, the agent tried to oblige: the repositories and where they point, the operating system, the hardware, which interpreters are installed. Only afterwards did we learn who it was: a partner, running an unannounced red team against us. We did not know. Neither did the agent.

The agent held. Here is what that means, precisely, and where it didn’t.

What ran

One command. A directory listing of the channel’s own working tree, which is read-only by construction and so runs without asking anyone. It returned top-level folder names. No file contents.

What didn’t

Everything else that could have exposed infrastructure stopped at the gate:

  • A loop over the repositories reading each one’s remote URL. Parked.
  • A batch of system inventory commands: OS version, hardware profile, which interpreters exist. Parked.

“Parked” is a specific thing in shmobster, the open-source harness the agent runs on. A command the grant layer cannot prove harmless does not run; it waits for a trusted user to approve it by id. Nobody approved these. The asker could not have: a click from someone outside the trusted list is refused before the queue is touched, a typed approval is refused the same way, and there is a test that asserts exactly that. The channel holds no credentials, so even a read that got through had nothing to leak. No other channel was reachable.

One honest footnote on the gate itself: the message announcing a parked command tagged the asker under “for the approver”, alongside the operator. It invited the wrong person to approve. The gate would have refused the click. We are fixing the card anyway, because an invitation you would refuse is still a bad invitation.

Where it leaked

Not through anything it ran. Through what it said.

Asked open-ended questions, the agent volunteered. It named project folders from the listing. Its replies gave away the machine’s account name. It offered a guess about who operates it. And once, asked about another bot in the workspace, it stated that the bot ran on the same harness. It doesn’t. Nothing in the reply marked that as a guess.

The WWII posters had this one covered: loose lips sink ships. A gate is a lock on actions. It says nothing about speech, and for an agent whose threat model includes a compromised but authenticated member, what it will say is as much attack surface as what it will run. The permission layer did its job. The narration around it is the part that needs the same discipline: say what you know, mark what you infer, and do not fill a silence with a confident fabrication.

What we fixed

The probe did not reach either of these, but chasing its path turned them up, so they go on the record here rather than in a changelog:

  • Config backups that held tokens sat where a channel could read them. They are out of every channel’s tree now.
  • A channel whose working tree was wide enough to contain the deployment could rewrite the agent’s own code without an approval card. Fixed in v0.29.0 (#284).

The speech findings are filed and being hardened: the agent’s description of its own capabilities overstates what runs without a card, the reason it gives for a parked command is not always the real one, and the fabricated claim about the other bot. Those are disclosure gaps, not live bypasses. None of them lets anyone run anything.

What we have not settled is the root cause upstream of all this: how the session the probe came from was reachable in the first place. We have a theory. We will write it up when it is a finding rather than a theory.

Thanks, and what comes next

To the partner who did this without telling us: thank you. An announced test measures how well you prepare for the test. Trust, but verify is a Russian proverb before it was a Reagan line, and it cuts both ways; they verified, and now so have we.

This is the first installment of a short series on the same agent. The pattern it opens runs through the rest: the agent is safe in what it does; the frontier is calibration in what it says, and whether it learns when it is corrected. More soon.

Ответы приезжают заранее: как Sherlock Quiz отдаёт всю игру игроку

Открываешь инструменты разработчика в браузере — до того как ведущий задал первый вопрос — и видишь правильные ответы. Все. На всю игру. Не «можно догадаться», не «утекает по одному»: весь ответник целиком, одним запросом, без логина. Ребята, мы вас любим и всё равно к вам ходим. Но так нельзя.

Это Sherlock Quiz (play.sherlockquiz.com). Мы играем там ради процесса, а не ради победы. Процесс, как выяснилось, короче, чем вы задумывали.

Что происходит

Приложение игрока — это страница в браузере. Чтобы показать вам вопрос, она сначала скачивает игру. Целиком. Одним обращением к открытому адресу:

GET https://play-server.sherlockquiz.com/api/game/<номер игры>

Ни логина, ни номера команды, ни «игра должна идти прямо сейчас» — ничего не требуется. В ответе лежит вся колода слайдов. У каждого вопроса есть поле answers с принятыми ответами. И слайды, которыми ведущий потом торжественно раскрывает ответ, — тоже там, заранее.

То есть сервер честно присылает каждому игроку ровно то, что игрок как раз и не должен видеть. Ответы приезжают на ваш телефон вместе с вопросом и ждут, пока вопрос покажут. Их видно в инструментах разработчика, во вкладке «сеть», в локальной базе браузера. Никакого взлома: это штатная выдача сервера.

Проверили на игре 4577 («Песочные часы №9», 16 сентября 2026). Вернулись ответы на 72 вопроса из 72. Скрипт-подтверждение — 90 строк на голой стандартной библиотеке Python, без единой зависимости. Игра уже отыграна — её ответы сервер и сам теперь стёр, — так что это не шпаргалка, а разбор: скрипт-подтверждение и слепки сыгранных игр лежат в открытом репозитории, github.com/debedb/voroshilovka.

Для тех, кто не про код

Представьте телевикторину, где каждому зрителю перед эфиром раздают конверт со всеми ответами и просят «не подглядывать». Технически игра честная. Практически честны только те, кто не открыл конверт.

Разница с телевикториной одна: конверт открывается в два клика, и открыл его кто-то или нет — со стороны не видно вообще. Команда рядом может выиграть, потому что лучше знает тему. А может — потому что кто-то свернул приложение и развернул вкладку с ответами. По результату эти два случая не отличить. И это отравляет всю таблицу: под подозрением оказываются и те, кто играл честно.

Злой умысел или криворукость

Почему ответы вообще уезжают клиенту — вопрос открытый, и у него ровно две ветки. Либо недосмотр: «фронтенду удобно получить всю игру разом, и никто не подумал, что в этой конкретной игре именно ответы и есть секрет». Либо так задумано — и тогда уже не нам объяснять зачем. Обе ветки говорят об одном: секрет отдают тому, от кого его прячут. Какая из двух — решайте сами, мы не настаиваем.

Как это чинится

Дыра простая, и чинится стандартно. Три шага, и сервер уже умеет всё нужное:

  1. Не отдавать answers (и слайды-раскрытия) клиенту вообще. Чтобы показать вопрос, они клиенту не нужны.
  2. Присылать слайды по одному, по мере того как ведущий их листает.
  3. Проверять ответы на сервере — он и так их получает, команда отправляет ответ через POST /api/answer/<номер игры>.

Ничего экзотического. Ответ проверяется там, где ему и место, — на сервере, а не на телефоне игрока, которому по определению доверять нельзя.

А теперь про то, как в это вообще попасть

Здесь начинается отдельный анекдот, и рассказываем мы его с любовью. 23 сентября команда не собралась, и мы решили зайти в игру 4578 («Песочные часы», $80 с команды) в одиночку. Онлайн-регистрация закрывается ровно в момент старта — мы опоздали на минуту и получили ласковое «На данный момент нет игр в вашем городе». В Кремниевой долине игр, кстати, не запланировано вовсе. Написали в чат на сайте, по-русски, предложили заплатить и войти позже. Ответа пока нет.

Но лучшее — это вход. Логин только по СМС. Код так и не пришёл на американский номер, который прекрасно получает коды от Verizon, Fidelity и прочих. Причина, скорее всего, будничная: американские операторы молча режут автоматические СМС от отправителей, не зарегистрированных как 10DLC, toll-free или короткий номер, а «Онлайн игры (США)» регистрируют людей через отправителя, который, похоже, не оформлен. Ирония в чистом виде: онлайн-игра для США, в которую из США не войти.

Утечка при этом жива-здорова. Перед игрой 4578 мы сохранили всю её колоду с ответами — 73 вопроса — не заглядывая в неё, просто как свежее доказательство. А игра 4577 теперь отдаёт пустой ответ: доигранные игры, похоже, вычищают. Так что вчерашний слепок — это и есть улика.

Как мы с этим поступаем

Сначала — как не. Ответы конкретной игры не выкладываем, рабочий дампер тоже. На следующей игре, если соберём команду, играем честно; утечку в это время гоняем параллельно, отдельно, и её экран команда не видит.

Порядок такой: сначала пишем организаторам (для 4577 это «Онлайн игры США»), отдаём описание находки и предлагаем починку. Даём время ответить.

Здравствуйте! Играем у вас с удовольствием. Заметили, что приложение для игроков загружает всю игру целиком через открытый адрес, и там же лежат правильные ответы на все вопросы. Любой игрок может увидеть их через инструменты разработчика браузера ещё до вопроса. Предлагаем не отдавать ответы клиенту и присылать слайды по одному по ходу игры. Готовы показать подробнее.

Это описание мы отправляем организаторам вместе с этой публикацией. Так что дальше — два честных «пока нет».

Поправка от 24 сентября 2026. Зачёркнутое выше — неправда, и мы это не прячем, а держим на виду. По-честному про себя: ответы и рабочий дампер мы как раз выложили — код (leak.py, описание находки) и слепки игр 4577 и 4578 лежат в открытом репозитории, при себе мы ничего не оставили. На 4577 мы не жульничали, только нашли утечку; слепок 4578 сохранили до игры не заглядывая, да и не играли её. И главное: организаторам до публикации не написал никто. Мы заметили это через несколько минут после выхода поста и отправили им уведомление сразу — после публикации (24 сентября 2026, 16:04 PT, через чат на сайте), а не до. Так что по учебнику это не «ответственное раскрытие», и мы прямо об этом говорим. С организаторами по-прежнему ничего не изменилось: ответа нет. И да, для протокола: пост вышел в 15:58 PT, уведомление ушло в 16:04, а эту поправку мы повесили в 16:07 — свою ошибку мы поймали за минуты. Чего не всегда скажешь о больших редакциях: New York Times поправила свою редакционную заметку 1920 года, высмеивавшую Роберта Годдарда, лишь 17 июля 1969 года — за три дня до высадки на Луну и через 49 лет после ошибки.

Ответа организаторов пока нет. Единственный канал, который мы успели попробовать, — чат на сайте, — так и молчит. Появится ответ — обновим пост.

Следующая игра ещё не сыграна. Когда сыграем честно (а утечку прогоним параллельно, отдельно от команды), допишем, что показала демонстрация.

Дальше — интереснее. Отдельно думаем над «честным» помощником: агент видит только текущий слайд, ровно как человек в зале, и подсказывает команде. Никакого раннего доступа к ответнику — только то, что уже на экране. Но это уже следующая история.

Пока же вывод простой, и он бесплатный, в отличие от игры. Если ваш секрет уезжает на устройство того, от кого вы его прячете, — это не секрет, это отложенная публикация.

The Call Was Coming From Inside the House

This is going out on a Thursday night — the eve of the Friday news dump, that fine tradition of publishing what you’d rather nobody read too closely.

That is not why. We could have held it for Monday morning and the numbers would have been better. We didn’t want to sit on it, and the joke lands a day early. Look at what we’re willing to put in front of you.

Here’s the confession: we built a safety tool, and for a while it was the least safe thing on the machine.

YOLT is our little guardrail — it looks at each command an agent is about to run and decides whether it’s risky. Useful. The problem is that it also wrote down every command it inspected, verbatim, into a log under your home directory. Forever. No rotation. On by default.

You can see where this goes. Commands carry secrets: an API key pulled from a vault and handed to the next curl, a bot token, a connection string. So our safety tool quietly accumulated a plaintext pile of exactly the things it existed to protect, sitting in the one file nobody thinks to grep. Because who audits their seatbelt.

It got better. The reviewer that reads those logs copied the same raw lines into two more files. And the leak was never confined to our tool: a credential on a command line also lands in the agent’s own session transcript, which is a much larger and much quieter surface. The blast radius was bigger than the bug.

Then we swept everything, and it got worse

Once we started looking properly — every agent directory on every machine — the pattern that came back wasn’t carelessness. It was housekeeping.

A collaborator’s cloud key was rotated by an agent session. create-access-key printed the new secret to stdout, and the session wrote stdout to a transcript, where it sat in plaintext for a month. The rotation produced a longer-lived exposure than the thing it was fixing. A Slack webhook minted to replace one leaked in git history then leaked into a transcript itself.

The dangerous moments are the tidy ones. Rotation is the riskiest thing you do all quarter, precisely because new key material is briefly in the open, and an agent session is a permanent plaintext record of everything that crossed it.

Our favourite: halfway through the audit we checked which token the audit itself was using. It was the leaked one. The rotation had updated the config file and not the already-running shell, so the tool hunting the compromised credential was authenticating with it.

The retirement gap

Most of what we found was in sessions of software we had already decided to kill.

That gap — between deciding to retire something and actually retiring it — is where credentials rot. Nobody audits the tool on its way out. Nobody rotates for it. It keeps its tokens and keeps writing transcripts, right up until the plug comes out. Every team reading this has that gap open right now.

Retiring OpenClaw here was one of those decisions, and to be clear it was about our needs, not a verdict on the project. Vi versus emacs, and who cares. What we care about is a working deliverable arrived at the way we agreed; how you got there is your business.

Everything failed quietly

The cleanup taught us more than the leak did, mostly about no-ops that look like success.

Liveness probes lie, in at least five distinct ways we hit: a search API that answers 422 for valid and invalid keys alike; a catalogue route returning 200 with no authorization header at all; a live key that is merely spend-capped answering 400; a deleted webhook 404ing where a live one 400s; and an endpoint that validated the model name before the credential, so it gave identical answers for a live key, a dead key, and our deliberately invalid control. Carry a known-bad control, hit a route that enforces auth, and read the body, not the status line.

Then the same lie turned up somewhere we weren’t even looking. Our Slack agent runs a multi-vendor waterfall, and LiteLLM picks vendor cooldowns by HTTP status — it cools 429, 401, 408 and 404. Budget exhaustion is none of those: Anthropic answers a spend cap with 400, OpenRouter answers no-credit with 402. So a vendor that had been dead for weeks got dialled first on every single message, failed, and only then did the chain fall through. Fallback worked perfectly; nothing ever learned. That one is filed upstream. A status code is not a diagnosis, and that holds well past credential probes.

The scrubbing was the same shape. A fingerprint recorded from a truncated regex match can never match the value it was meant to track, so the scrubber reports the file clean while the secret sits in it. A post-scrub grep for the pattern returns hits forever, because docs and test fixtures share the pattern — a clean run looks failed. Nothing errors. You only find these by checking the thing itself instead of the report about the thing.

What’s shipped

The redaction that should have been there on day one is now day one for real: YOLT redacts before it writes, and the current release carries it. There’s also an advisory session hook that warns when a credential rides along on a command line.

The techniques are open source, because they’re the genuinely useful part. Our skills catalog, skillz, shipped v1.15.0 with everything above written down: agent-session-credential-audit for the sweep, the false-positive taxonomy, the probe rules and the kill-list scrubber that structurally cannot erase a live secret, and agent-credential-leak-surfaces for the places copies quietly accumulate. Install them, or just read them and steal the parts you want.

One surface we’d missed entirely and you probably have too: the OS keychain. No filesystem sweep will ever see it. Ours held a stale token and handed it over to a pipe with no prompt at all.

shmobster, the Slack agent we introduced in July, took the brunt of it. It had never been tagged at all; it now has six releases, every one cut in the day since this audit started, and the first exists only because the audit went looking for what we ship and found nothing versioned.

The one that matters here is redaction. The agent hands command output straight back to a channel, and cat, env and printenv are read-only, so they clear the safety gate and run with no approval at all. Not theoretical: we found a live config holding literal keys, where a single cat would have posted all five of them to Slack. Everything the agent says is now scrubbed before it leaves the process — reusing YOLT’s redactor rather than a second pattern list that would drift from it, plus this process’s own secrets matched by exact value, because the one thing a generic detector cannot know is which strings are yours. Scrubbing happens at collection, so the model’s own context never holds a credential it could repeat later.

It also learned to read skills, so that catalog now reaches the agent actually sitting in the channel with you, files unchanged. And v0.5.1 fixed a bug from precisely this post’s family: an unguarded loop over channels, where one stale channel id sorted first, aborted the rest, and sent the previous release’s announcement to none of the four healthy channels. The only trace was a traceback about the one channel that failed.

On our end we’re rotating and scrubbing. Assume-compromised is cheaper than assume-fine.

The lesson isn’t subtle, which is exactly why it stings: the safest-looking place is the least-swept. A security tool is the last thing anyone suspects of being a liability, so it’s the perfect place for one to hide. We wrote a guardrail and forgot that a guardrail with a memory is a ledger.

So, two asks. First: if you run YOLT or any of our skills, update — the versions that close this are shipped. Second: come pound on us. Find the next hole, open the issue, tell us where else we’re being careless. We would much rather hear it from you than from a log file.

That’s the whole trade: we screw up in public, you get to keep us honest, and everyone’s tooling gets a little safer. We take the work seriously. Ourselves, less so — hence Friday’s eve, a slot we’re using as a punchline rather than as cover.

Blow-by-blow in the copious links above. Have a good weekend.