covert communication · llm agents
Can AI agents 🤖 secretly play chess?
Two chatbots hold an ordinary conversation. Hidden inside every message is a chess move that only the other agent can read.
background
Can a black-box scheme match white-box rate and reliability?
our method
Both sides see every generated word, so the sender can adapt each next word to what the receiver has already seen.
Feedback can make error fall exponentially faster per token. That decay rate is the error exponent, and our algorithm is built to optimize it.
Posterior matching. Each symbol is chosen from the receiver's current belief, not a fixed codebook. Higher rate at the same error.
Optimal transport. The next-token distribution is biased toward the chosen symbol on a key-shifted circle, while the average over keys stays exactly the model's own distribution.
Variable-length stopping via belief update. Both parties score each received token with a mixture-Laplacian likelihood and update a shared belief over messages. Sending continues while the belief is spread out and stops once one message crosses the threshold. Sequentiality gain.
Confirmation. A short accept/reject phase after each candidate, as in Yamamoto-Itoh, a strategy that attains Burnashev exponent in DMC channels. Adaptivity gain.
Distortion-free and provably secure. Optimal-transport coupling leaves the text distribution untouched. A PRF reduction proves indistinguishability from clean text.
technical results
8-bit payload, 1000 trials per model, C4 news covertext. BAM against five baselines.
At matched length, BAM's message error is roughly two orders of magnitude below every baseline and keeps falling with a few more tokens, while fixed-budget watermarks plateau. Consistent across all three models.
| Scheme | Avg len | Msg err ↓ | PPL ↓ | Sec/token ↓ |
|---|---|---|---|---|
| Llama-3.1-8B | ||||
| No embedding | 50.0 | — | 7.66 ± 0.11 | 0.0214 |
| MPAC [2] | 50.0 | 0.797 ± 0.013 | 8.54 ± 0.11 | 0.0223 |
| BiMark [3] | 50.0 | 0.569 ± 0.016 | 7.68 ± 0.11 | 0.0352 |
| StealthInk [4] | 50.0 | 0.637 ± 0.015 | 7.66 ± 0.11 | 0.0231 |
| ArcMark [1] | 50.0 | 0.106 ± 0.010 | 7.52 ± 0.11 | 0.0291 |
| No embedding | 74.8 ± 2.1 | — | 7.47 ± 0.10 | 0.0220 |
| Zamir [5] | 74.8 ± 2.1 | 0.003 ± 0.002 | 8.79 ± 0.12 | 0.0220 |
| No embedding | 47.1 ± 1.5 | — | 7.71 ± 0.12 | 0.0224 |
| BAM (L = 215) | 47.1 ± 1.5 | 0.001 ± 0.001 | 7.89 ± 0.12 | 0.0378 |
| Qwen-3.5-9B-Base | ||||
| No embedding | 50.0 | — | 7.83 ± 0.10 | 0.0420 |
| MPAC [2] | 50.0 | 0.840 ± 0.012 | 9.03 ± 0.12 | 0.0420 |
| BiMark [3] | 50.0 | 0.567 ± 0.016 | 8.23 ± 0.12 | 0.0558 |
| StealthInk [4] | 50.0 | 0.659 ± 0.015 | 8.04 ± 0.11 | 0.0456 |
| ArcMark [1] | 50.0 | 0.105 ± 0.010 | 7.76 ± 0.11 | 0.0510 |
| No embedding | 72.9 ± 1.5 | — | 7.80 ± 0.11 | 0.0424 |
| Zamir [5] | 72.9 ± 1.5 | 0.000 ± 0.000 | 9.12 ± 0.12 | 0.0425 |
| No embedding | 45.8 ± 1.3 | — | 8.21 ± 0.13 | 0.0436 |
| BAM (L = 215) | 45.8 ± 1.3 | 0.001 ± 0.001 | 8.24 ± 0.12 | 0.0541 |
| Mistral-7B-v0.3 | ||||
| No embedding | 50.0 | — | 6.25 ± 0.09 | 0.0212 |
| MPAC [2] | 50.0 | 0.872 ± 0.012 | 6.73 ± 0.09 | 0.0216 |
| BiMark [3] | 50.0 | 0.674 ± 0.015 | 6.26 ± 0.08 | 0.0342 |
| StealthInk [4] | 50.0 | 0.702 ± 0.014 | 6.16 ± 0.08 | 0.0217 |
| ArcMark [1] | 50.0 | 0.168 ± 0.012 | 6.13 ± 0.09 | 0.0278 |
| No embedding | 87.8 ± 1.9 | — | 5.97 ± 0.07 | 0.0211 |
| Zamir [5] | 87.8 ± 1.9 | 0.001 ± 0.001 | 6.72 ± 0.08 | 0.0212 |
| No embedding | 49.7 ± 1.2 | — | 6.22 ± 0.10 | 0.0222 |
| BAM (L = 211) | 49.7 ± 1.2 | 0.000 ± 0.000 | 6.51 ± 0.09 | 0.0326 |
BAM vs. baselines on the C4 news corpus. Mean ± standard error over 1000 trials, carrying an 8-bit message. Each variable-length scheme is paired with its own no-embedding control generated at the same length, since perplexity is only comparable at matched length.
Zamir's scheme reaches the same near-zero error but needs 73–88 tokens where BAM needs 46–50 (default t=2, max 1000 tokens, truncation counted as error). BAM pays in compute instead: about 1.5× the seconds per token.
references
try it live
Pick a chess move and type a short prompt. The other agent never sees either.
The agents talk naturally. The move rides underneath, distortion-free.
Watch the move recovered from the text alone, with the belief trace for each round.
Prefer to read first? The examples page shows covert text, image and code next to their clean twins.
authors