📌 이 취약점에 대해 확인된 사실
전부 발행처가 발표한 값입니다. 우리가 계산하거나 판단한 숫자는 하나도 없습니다.
악용 여부
심각도 (발행처 발표값)
높음7.1
CVSS:4.0/AV:N/AC:L/AT:N/PR:L/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N/E:X/CR:X/IR:X/AR:X/MAV:X/MAC:X/MAT:X/MPR:X/MUI:X/MVC:X/MVI:X/MVA:X/MSC:X/MSI:X/MSA:X/S:X/AU:X/R:X/V:X/RE:X/U:X
악용 확률 (EPSS)
0.3%
📄 원문 그대로
아래 문장은 전부 발행처가 쓴 것입니다. 번역하지 않습니다 — 보안 문서의 오역은 조치를 바꿉니다.
취약점 설명 (NVD)
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimodal EngineInput directly from the caller-supplied token_ids, and GenerateRequest.token_ids (vllm/entrypoints/serve/disagg/protocol.py) is not checked against model_config.max_model_len. For multimodal processors that report skip_prompt_length_check=True (for example Nemotron Parse, Whisper, and FireRedLID), InputProcessor._validate_prompt_len() returns immediately for both encoder and decoder prompts, so an overlong prompt becomes an EngineCoreRequest and reaches the worker input-batch copy into a fixed max_model_len-wide NumPy row. A client able to reach the endpoint on an affected model configuration can therefore submit an overlong token_ids list to trigger a worker failure and denial of service. Fixed in 0.29.0.
참고
악용 확률 변화
우리가 매일 저장한 EPSS 스냅샷입니다. 원본은 전날 값만 주므로, 이 표는 수집을 시작한 이후만 보여줍니다.
| 기준일 | 확률 | 백분위 |
|---|---|---|
| 2026-10-02 | 0.31% | 21.3% |
| 2026-10-01 | 0.31% | 21.2% |
| 2026-09-29 | 0.31% | 21.1% |
| 2026-09-27 | 0.31% | 21.0% |
🧩 같은 약점 유형 — CWE-400
같은 분류의 다른 취약점입니다. 같은 실수가 제품을 가리지 않고 반복됩니다.