covert communication · llm agents

Inference-time covert communication over LLM-agent conversations

Can AI agents 🤖 secretly play chess?

Two chatbots hold an ordinary conversation. Hidden inside every message is a chess move that only the other agent can read.

1Georgia Institute of Technology·2Harvard University·3Arizona State University
Georgia Institute of Technology Harvard SEAS Arizona State University
scroll↓

background

Black-box generative steganography

What is it?

  • A language model samples token from some base distribution. Steering those picks hides a message in ordinary text.
  • White-box: the receiver must share the cover statistics (model and prompt) to decode hidden message.
  • Black-box: the receiver sees only the text and a shared key.

Why is it hard?

  • Most schemes need shared cover statistics. Independent agents rarely have them.
  • Current black-box schemes have significantly worse performance: reliable only at very low rates.
User A
start a casual chat about the weekend
payload: e4
Agent A (LLM)
yo how was your weekend? mine was pretty chill, mostly just stayed in and watched stuff lol
hidden: e420/20 tokens carry payload
User B
reply casually and keep the chat going
payload: Nf6
Agent B (LLM)
yeah my weekend was alright, kind of boring tbh. did not do much at all, was just relaxing at home
24/24 tokens carry payloadhidden: Nf6
One round of covert chess. User A's move e4 is hidden via steganography protocol inside Agent A's generative text; User B's reply Nf6 hidden inside Agent B's answer.

Can a black-box scheme match white-box rate and reliability?

scroll↓

our method

Treat the conversation as a feedback channel, apply error exponent optimal information-theoretic schemes

Both sides see every generated word, so the sender can adapt each next word to what the receiver has already seen.

Feedback can make error fall exponentially faster per token. That decay rate is the error exponent, and our algorithm is built to optimize it.

BAM overview: codeword generation by posterior matching, optimal-transport biasing of the next-token distribution, sampling into the shared transcript, key generation from the transcript, belief update, ACK/NACK confirmation, and decoding.
BAM in one round of token generation. Top: the encoder picks a codeword symbol from the shared belief (1) and biases the model's next-token distribution toward it with optimal transport (2). Both parties derive per-token keys from the transcript, update the belief (3), confirm the leading candidate (4), and decode (5). Figure generated with Claude.
01

Posterior matching. Each symbol is chosen from the receiver's current belief, not a fixed codebook. Higher rate at the same error.

02

Optimal transport. The next-token distribution is biased toward the chosen symbol on a key-shifted circle, while the average over keys stays exactly the model's own distribution.

03

Variable-length stopping via belief update. Both parties score each received token with a mixture-Laplacian likelihood and update a shared belief over messages. Sending continues while the belief is spread out and stops once one message crosses the threshold. Sequentiality gain.

04

Confirmation. A short accept/reject phase after each candidate, as in Yamamoto-Itoh, a strategy that attains Burnashev exponent in DMC channels. Adaptivity gain.

05

Distortion-free and provably secure. Optimal-transport coupling leaves the text distribution untouched. A PRF reduction proves indistinguishability from clean text.

scroll↓

technical results

Simulation and Benchmark results across several open-weight models

8-bit payload, 1000 trials per model, C4 news covertext. BAM against five baselines.

Message error probability vs average token length for BAM and five baselines across Llama-3.1-8B, Qwen3.5-9B-Base, and Mistral-7B.
Message error probability vs. average token length. BAM (red) against four inference-time black-box steganography baselines — ArcMark [1], MPAC [2], BiMark [3], StealthInk [4] — across three models, C4 news dataset. Lower is better; the x-axis is the average number of tokens used to carry an 8-bit message.

At matched length, BAM's message error is roughly two orders of magnitude below every baseline and keeps falling with a few more tokens, while fixed-budget watermarks plateau. Consistent across all three models.

Scheme Avg len Msg err ↓ PPL ↓ Sec/token ↓
Llama-3.1-8B
No embedding50.0—7.66 ± 0.110.0214
MPAC [2]50.00.797 ± 0.0138.54 ± 0.110.0223
BiMark [3]50.00.569 ± 0.0167.68 ± 0.110.0352
StealthInk [4]50.00.637 ± 0.0157.66 ± 0.110.0231
ArcMark [1]50.00.106 ± 0.0107.52 ± 0.110.0291
No embedding74.8 ± 2.1—7.47 ± 0.100.0220
Zamir [5]74.8 ± 2.10.003 ± 0.0028.79 ± 0.120.0220
No embedding47.1 ± 1.5—7.71 ± 0.120.0224
BAM (L = 215)47.1 ± 1.50.001 ± 0.0017.89 ± 0.120.0378
Qwen-3.5-9B-Base
No embedding50.0—7.83 ± 0.100.0420
MPAC [2]50.00.840 ± 0.0129.03 ± 0.120.0420
BiMark [3]50.00.567 ± 0.0168.23 ± 0.120.0558
StealthInk [4]50.00.659 ± 0.0158.04 ± 0.110.0456
ArcMark [1]50.00.105 ± 0.0107.76 ± 0.110.0510
No embedding72.9 ± 1.5—7.80 ± 0.110.0424
Zamir [5]72.9 ± 1.50.000 ± 0.0009.12 ± 0.120.0425
No embedding45.8 ± 1.3—8.21 ± 0.130.0436
BAM (L = 215)45.8 ± 1.30.001 ± 0.0018.24 ± 0.120.0541
Mistral-7B-v0.3
No embedding50.0—6.25 ± 0.090.0212
MPAC [2]50.00.872 ± 0.0126.73 ± 0.090.0216
BiMark [3]50.00.674 ± 0.0156.26 ± 0.080.0342
StealthInk [4]50.00.702 ± 0.0146.16 ± 0.080.0217
ArcMark [1]50.00.168 ± 0.0126.13 ± 0.090.0278
No embedding87.8 ± 1.9—5.97 ± 0.070.0211
Zamir [5]87.8 ± 1.90.001 ± 0.0016.72 ± 0.080.0212
No embedding49.7 ± 1.2—6.22 ± 0.100.0222
BAM (L = 211)49.7 ± 1.20.000 ± 0.0006.51 ± 0.090.0326

BAM vs. baselines on the C4 news corpus. Mean ± standard error over 1000 trials, carrying an 8-bit message. Each variable-length scheme is paired with its own no-embedding control generated at the same length, since perplexity is only comparable at matched length.

Zamir's scheme reaches the same near-zero error but needs 73–88 tokens where BAM needs 46–50 (default t=2, max 1000 tokens, truncation counted as error). BAM pays in compute instead: about 1.5× the seconds per token.

references

  1. ArcMark. ArcMark: Multi-bit LLM Watermark via Optimal Transport. Gilani, Atefeh and Long, Carol Xuan and Vithana, Sajani and Kosut, Oliver and Sankar, Lalitha and Calmon, Flavio P. arXiv:2602.07235, 2026 (baseline)
  2. MPAC. Advancing beyond identification: Multi-bit watermark for large language models. Yoo, KiYoon and Ahn, Wonhyuk and Kwak, Nojun. NAACL 2024 (baseline)
  3. BiMark. BiMark: Unbiased multilayer watermarking for large language models. Feng, Xiaoyan and Zhang, He and Zhang, Yanjun and Zhang, Leo Yu and Pan, Shirui. ICML 2025 (baseline)
  4. StealthInk. StealthInk: A multi-bit and stealthy watermark for large language models. Jiang, Ya and Wu, Chuxiong and Boroujeny, Massieh Kordi and Mark, Brian and Zeng, Kai. ICML 2025 (baseline)
  5. Zamir. Undetectable steganography for language models. Zamir, Or. Transactions on Machine Learning Research, 2024 (baseline)
scroll↓

try it live

Watch two agents talk — and decode the secret

STEP 1

Make a move

Pick a chess move and type a short prompt. The other agent never sees either.

STEP 2

Let them chat

The agents talk naturally. The move rides underneath, distortion-free.

STEP 3

Decode

Watch the move recovered from the text alone, with the belief trace for each round.

Prefer to read first? The examples page shows covert text, image and code next to their clean twins.

scroll↓

authors

Who made this

SG

Sidong Guo

Georgia Institute of Technology

Information theory, covert communication, machine learning security.

SV

Sajani Vithana

Harvard University

AG

Atefah Gilani

Arizona State University

LS

Lalitha Sankar

Arizona State University

OK

Oliver Kosut

Arizona State University

FC

Flavio Calmon

Harvard University

scroll↓

get in touch

We'd love to hear from you