0.8.4
This release contains 180 commits from 84 contributors (25 new contributors!).
Highlights
This release includes important accuracy fixes for Llama4 models, if you are using it, we highly recommend you to update.
Model
- Llama4 (#16113,#16509) bug fix and enhancements:
- qknorm should be not shared across head (#16311)
- Enable attention temperature tuning by default for long context (>32k) (#16439)
- Index Error When Single Request Near Max Context (#16209)
- Add tuned FusedMoE kernel config for Llama4 Scout, TP=8 on H100 (#16488)
- Update to transformers==4.51.1 (#16257)
- Added chat templates for LLaMa4 pythonic tool calling (#16463)
- Optimized topk for topk=1(#16512)
- Add warning for Attention backends that do not support irope yet (#16212)
- Support Qwen3 and Qwen3MoE (#15289), smolvlm (#16017), jinaai/jina-embeddings-v3 (#16120), InternVL3 (#16495), GLM-4-0414 (#16338)
API
- Estimate max-model-len use available KV cache memory. The error message nows hints at how to set
--max-model-len(#16168) - Add hf_token to EngineArgs (#16093)
- Enable regex support with xgrammar in V0 engine (#13228)
- Support matryoshka representation / support embedding API dimensions (#16331)
- Add bucket for
request_latency,time_to_first_tokenandtime_per_output_token(#15202) - Support for TorchAO quantization (#14231)
Hardware
- Intel-Gaudi: Multi-step scheduling implementation for HPU (#12779)
- TPU:
- Make @support_torch_compile work for XLA backend (#15782)
- Use
language_modelinterface for getting text backbone in MM (#16410)
Performance
- DeepSeek MLA: a new merge_attn_states CUDA kernel, 3x speedup (#16173)
- MoE: Support W8A8 channel-wise weights and per-token activations in triton fused_moe_kernel (#16366)
- Add support to modelopt quantization of Mixtral model (#15961)
- Enable PTPC FP8 for CompressedTensorsW8A8Fp8MoEMethod (triton fused_moe) (#16537)
V1 Engine Core
- Enable multi-input by default (#15799)
- Scatter and gather placeholders in the model runner (#16076)
- Set structured output backend to
autoby default (#15724) - Zero-copy tensor/ndarray serialization/transmission (#13790)
- Eagle Model loading (#16035)
- KV cache slots for eagle heads (#16370)
- Add
supports_structured_output()method to Platform (#16148)
Developer Facing
- Add sampling parameters to benchmark_serving. (#16022)
- AutoWeightsLoader refacotring (#16383, #16325, #16088, #16203, #16103)
- Unifieid configuration with engine args:
LoadConfig(#16422),ParallelConfig(#16332)
What's Changed
- [Misc] Auto detect bitsandbytes pre-quantized models by @tristanleclercq in https://github.com/vllm-project/vllm/pull/16027
- [CI] Fix benchmark script level by @khluu in https://github.com/vllm-project/vllm/pull/16089
- fix: support clang17 for macos and fix the real libomp by @yihong0618 in https://github.com/vllm-project/vllm/pull/16086
- [doc] fix 404 by @reidliu41 in https://github.com/vllm-project/vllm/pull/16082
- Revert "doc: add info for macos clang errors (#16049)" by @yihong0618 in https://github.com/vllm-project/vllm/pull/16091
- Fix some capitalisations in generated examples doc titles by @hmellor in https://github.com/vllm-project/vllm/pull/16094
- [Misc] format output for encoder_decoder.py by @reidliu41 in https://github.com/vllm-project/vllm/pull/16095
- [Misc] Remove redundant code by @chaunceyjiang in https://github.com/vllm-project/vllm/pull/16098
- [Bugfix] fix use_atomic_add support of marlin kernel when using v1 engine by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/15946
- [Model] use AutoWeightsLoader for phi, gemma, deepseek by @jonghyunchoe in https://github.com/vllm-project/vllm/pull/16088
- [Model] fix model testing for TeleChat2ForCausalLM and V0 llama4 by @luccafong in https://github.com/vllm-project/vllm/pull/16112
- [Benchmark] Add sampling parameters to benchmark_serving. by @hyeygit in https://github.com/vllm-project/vllm/pull/16022
- [Frontend] Fix typo in tool chat templates for llama3.2 and toolace by @bjj in https://github.com/vllm-project/vllm/pull/14501
- [CI][V1] Fix passing
tokenizeras kwarg tovalidate_guidance_grammarby @ywang96 in https://github.com/vllm-project/vllm/pull/16117 - [Misc] refactor example eagle by @reidliu41 in https://github.com/vllm-project/vllm/pull/16100
- [Doc][Bugfix] Add missing EOF in k8s deploy doc by @psschwei in https://github.com/vllm-project/vllm/pull/16025
- [Misc] Improve model redirect to accept json dictionary by @Isotr0py in https://github.com/vllm-project/vllm/pull/16119
- [Model] use AutoWeightsLoader for stablelm,starcoder2,zamba2 by @lengrongfu in https://github.com/vllm-project/vllm/pull/16103
- [Bugfix] LoRA : Fix the order in which the kernels process LoRAs by @varun-sundar-rabindranath in https://github.com/vllm-project/vllm/pull/16040
- [Bugfix] add hf_token to EngineArgs by @paolovic in https://github.com/vllm-project/vllm/pull/16093
- [Misc] update requires-python in pyproject.toml by @reidliu41 in https://github.com/vllm-project/vllm/pull/16116
- [TPU] Update PyTorch/XLA by @yaochengji in https://github.com/vllm-project/vllm/pull/16130
- [V1][Minor] Optimize get_cached_block by @WoosukKwon in https://github.com/vllm-project/vllm/pull/16135
- Fix requires-python by @martinhoyer in https://github.com/vllm-project/vllm/pull/16132
- [Metrics] Add bucket for
request_latency,time_to_first_tokenandtime_per_output_tokenby @yankay in https://github.com/vllm-project/vllm/pull/15202 - [V1][Minor] Minor simplification for get_computed_blocks by @WoosukKwon in https://github.com/vllm-project/vllm/pull/16139
- [Misc] Update Mistral-3.1 example by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16147
- [Bugfix] Make dummy encoder prompt padding alternative and add missing warnings by @Isotr0py in https://github.com/vllm-project/vllm/pull/16129
- [CI] Set max transformers version for Ultravox model test by @ywang96 in https://github.com/vllm-project/vllm/pull/16149
- doc: fix some typos in doc by @yihong0618 in https://github.com/vllm-project/vllm/pull/16154
- [VLM] Florence-2 supports online serving by @Isotr0py in https://github.com/vllm-project/vllm/pull/16164
- [V1][Structured Output] Add
supports_structured_output()method to Platform by @shen-shanshan in https://github.com/vllm-project/vllm/pull/16148 - [Model] Add Qwen3 and Qwen3MoE by @YamPengLi in https://github.com/vllm-project/vllm/pull/15289
- [Misc] improve example mlpspeculator and llm_engine_example by @reidliu41 in https://github.com/vllm-project/vllm/pull/16175
- [Doc]Update image to latest version by @WangErXiao in https://github.com/vllm-project/vllm/pull/16186
- Upstream Llama4 Support to Main by @houseroad in https://github.com/vllm-project/vllm/pull/16113
- [Bugfix] Re-enable support for
ChatGLMForConditionalGenerationby @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16187 - [V1] Revert the default
max_num_seqsto V0 values for most hardware by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16158 - [Misc] Print encoder seq len to short warning only once by @gshtras in https://github.com/vllm-project/vllm/pull/16193
- [Misc] Human-readable
max-model-lencli arg by @NickLucche in https://github.com/vllm-project/vllm/pull/16181 - [Misc] Move Llama 4 projector call into encoder execution by @ywang96 in https://github.com/vllm-project/vllm/pull/16201
- [Bugfix] Fix guidance backend for Qwen models by @benchislett in https://github.com/vllm-project/vllm/pull/16210
- [V1][BugFix] Exit properly if engine core fails during startup by @njhill in https://github.com/vllm-project/vllm/pull/16137
- [Misc] add description attribute in CLI by @reidliu41 in https://github.com/vllm-project/vllm/pull/15921
- [Bugfix][V0] XGrammar structured output supports Enum by @leon-seidel in https://github.com/vllm-project/vllm/pull/15878
- Torchao by @drisspg in https://github.com/vllm-project/vllm/pull/14231
- [ROCm][Bugfix][FP8] Make fp8 quant respect fused modules mapping by @mgoin in https://github.com/vllm-project/vllm/pull/16031
- [core] do not send error across process by @youkaichao in https://github.com/vllm-project/vllm/pull/16174
- [Misc] Update compressed-tensors to version 0.9.3 by @mlsw in https://github.com/vllm-project/vllm/pull/16196
- Update BASE_IMAGE to 2.22 release of Neuron by @aws-satyajith in https://github.com/vllm-project/vllm/pull/16218
- [V1] Scatter and gather placeholders in the model runner by @ywang96 in https://github.com/vllm-project/vllm/pull/16076
- [Bugfix] fix use-ep bug to enable ep by dp/tp size > 1 by @zxfan-cpu in https://github.com/vllm-project/vllm/pull/16161
- Add warning for Attention backends that do not support irope yet by @sarckk in https://github.com/vllm-project/vllm/pull/16212
- [Bugfix] Do not skip "empty" parts of chats that are parsable by @mgoin in https://github.com/vllm-project/vllm/pull/16219
- [Bugfix] Fix and reorganize broken GGUF tests and bump gguf version by @Isotr0py in https://github.com/vllm-project/vllm/pull/16194
- [torch.compile][TPU] Make @support_torch_compile work for XLA backend by @lsy323 in https://github.com/vllm-project/vllm/pull/15782
- [V1] Add
disable_chunked_mm_inputarg to disable partial mm input prefill by @mgoin in https://github.com/vllm-project/vllm/pull/15837 - [Misc] Merge the logs of pp layers partitions by @kebe7jun in https://github.com/vllm-project/vllm/pull/16225
- [Docs] Add Slides from Singapore Meetup by @simon-mo in https://github.com/vllm-project/vllm/pull/16213
- [Misc] format and refactor some examples by @reidliu41 in https://github.com/vllm-project/vllm/pull/16252
- [Misc] Add warning for multimodal data in LLM.beam_search by @alex-jw-brooks in https://github.com/vllm-project/vllm/pull/16241
- [Model] use AutoWeightsLoader for phimoe,qwen2_moe,qwen3_moe by @lengrongfu in https://github.com/vllm-project/vllm/pull/16203
- [BugFix][ROCm] Fix GGUF MoE Dispatch Block_Dim for ROCm by @tywuAMD in https://github.com/vllm-project/vllm/pull/16247
- [Bugfix] Remove triton do_bench fast_flush arg by @kebe7jun in https://github.com/vllm-project/vllm/pull/16256
- Update to transformers==4.51.1 by @hmellor in https://github.com/vllm-project/vllm/pull/16257
- [New Model]: jinaai/jina-embeddings-v3 by @noooop in https://github.com/vllm-project/vllm/pull/16120
- [Misc] Avoid stripping meaningful whitespace from
nvidia-smi topo -moutput in collect_env.py by @imkero in https://github.com/vllm-project/vllm/pull/16272 - [Bugfix] Proper input validation for multi-modal encoder-decoder models by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16156
- [Bugfix] Handle
process_weights_after_loadingforQKVCrossParallelLinearby @Isotr0py in https://github.com/vllm-project/vllm/pull/15328 - Add warning that content below line in template will be removed by @hmellor in https://github.com/vllm-project/vllm/pull/16276
- [BugFix] Fix Llama4 - Index Error When Single Request Near Max Context by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/16209
- [Bugfix] fix deepseek fp16 scale bug by @jinzhen-lin in https://github.com/vllm-project/vllm/pull/14809
- [V1] Update structured output offline inference example by @russellb in https://github.com/vllm-project/vllm/pull/15721
- [CI/Build] Fix CI LoRA failure by @jeejeelee in https://github.com/vllm-project/vllm/pull/16270
- Add support to modelopt quantization of Mixtral model by @yueshen2016 in https://github.com/vllm-project/vllm/pull/15961
- [Model] Add smolvlm support by @chaunceyjiang in https://github.com/vllm-project/vllm/pull/16017
- [Bug] [ROCm] Fix Llama 4 Enablement Bug on ROCm: V0 ROCmFlashAttentionImpl and Triton Fused MoE bugs by @tjtanaa in https://github.com/vllm-project/vllm/pull/16198
- [Bugfix] fix gettid method is not define by @lengrongfu in https://github.com/vllm-project/vllm/pull/16084
- [Feature] Estimate max-model-len use available KV cache memory by @lengrongfu in https://github.com/vllm-project/vllm/pull/16168
- [Core] Upgrade to xgrammar 0.1.18, add cache size limit by @russellb in https://github.com/vllm-project/vllm/pull/16283
- [CI][Bugfix] Fix bad tolerance for test_batch_base64_embedding by @mgoin in https://github.com/vllm-project/vllm/pull/16221
- [TPU] Update PyTorch/XLA by @yaochengji in https://github.com/vllm-project/vllm/pull/16288
- [BugFix] Fix fusion test and add them to CI by @ProExpertProg in https://github.com/vllm-project/vllm/pull/16287
- [Misc] Fix test_sharded_state_loader.py(#16004) by @Accelerator1996 in https://github.com/vllm-project/vllm/pull/16005
- [Bugfix] Avoid transferring cached multi-modal items from P0 to P1 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16273
- Update label-tpu mergify and remove removal bot by @mgoin in https://github.com/vllm-project/vllm/pull/16298
- [BugFix] logger is not callable by @yihong0618 in https://github.com/vllm-project/vllm/pull/16312
- [BugFix] llama4 qknorm should be not shared across head by @luccafong in https://github.com/vllm-project/vllm/pull/16311
- update neuron config by @ajayvohra2005 in https://github.com/vllm-project/vllm/pull/16289
- [BugFix] fix some typos found by typos. by @yihong0618 in https://github.com/vllm-project/vllm/pull/16314
- [Model] Add
SupportsMultiModal.get_language_modelinterface by @NickLucche in https://github.com/vllm-project/vllm/pull/16007 - [Bugfix][Frontend] respect provided default guided decoding backend by @gcalmettes in https://github.com/vllm-project/vllm/pull/15476
- Revert "Update label-tpu mergify and remove removal bot" by @mgoin in https://github.com/vllm-project/vllm/pull/16350
- [Bugfix] Fix profiling.py by @hhy3 in https://github.com/vllm-project/vllm/pull/16202
- [Bugfix] catch AssertionError in MistralTokenizer as ValueError by @gcalmettes in https://github.com/vllm-project/vllm/pull/16344
- [CI]Fix hpu docker and numpy version for CI by @xuechendi in https://github.com/vllm-project/vllm/pull/16355
- Fix
benchmark_throughput.py --backend=hfby @mgoin in https://github.com/vllm-project/vllm/pull/16352 - [Build/CI] Add tracing deps to vllm container image by @russellb in https://github.com/vllm-project/vllm/pull/15224
- [Hardware] add platform-specific request validation api by @joerunde in https://github.com/vllm-project/vllm/pull/16291
- [Misc] refactor Structured Outputs example by @reidliu41 in https://github.com/vllm-project/vllm/pull/16322
- [TPU][V1] Refine tpu_model_runner to mitigate future recompilation issues by @yaochengji in https://github.com/vllm-project/vllm/pull/16275
- Add GLM-4-0414 support by @zRzRzRzRzRzRzR in https://github.com/vllm-project/vllm/pull/16338
- [Bugfix]: do not shutdown server if
skip_special_use=Falsefor MistralTokenizer by @gcalmettes in https://github.com/vllm-project/vllm/pull/14094 - [Model] use AutoWeightsLoader for granite, granitemoe, granitemoeshared, grok1, mixtral by @aaron-ang in https://github.com/vllm-project/vllm/pull/16325
- [TPU] Fix dummy loading OOM by @yaochengji in https://github.com/vllm-project/vllm/pull/16372
- [bugfix] Avoid the time consumption caused by creating dummy videos. by @Jintao-Huang in https://github.com/vllm-project/vllm/pull/16371
- [CI][Bugfix] Pin triton version for CPU by @ywang96 in https://github.com/vllm-project/vllm/pull/16384
- [misc] use tqdm.auto where appropriate by @BKitor in https://github.com/vllm-project/vllm/pull/16290
- [Bugfix][TPU] Fix TPU validate_request by @mgoin in https://github.com/vllm-project/vllm/pull/16369
- fix sonnet dataset sample when prefix len is very small by @Chenyaaang in https://github.com/vllm-project/vllm/pull/16379
- [Model] use AutoWeightsLoader for deepseek_v2, internlm2 by @aaron-ang in https://github.com/vllm-project/vllm/pull/16383
- [Misc] Update transformers version limits of multi-modal tests by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16381
- [Bugfix] Fix validation error for text-only Mllama 3.2 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16377
- [Kernel] Use moe_wna16 kernel for compressed tensors wna16 moe models by @mgoin in https://github.com/vllm-project/vllm/pull/16038
- [doc] add download model tips by @reidliu41 in https://github.com/vllm-project/vllm/pull/16389
- Update Numba to 0.61.2 by @cyyever in https://github.com/vllm-project/vllm/pull/16376
- [Model] Remove image mm limit for LLaMa4 by @yeqcharlotte in https://github.com/vllm-project/vllm/pull/16365
- [doc] update the wrong link by @reidliu41 in https://github.com/vllm-project/vllm/pull/16401
- [CI] Add auto update workflow for Dockerfile graph by @WineChord in https://github.com/vllm-project/vllm/pull/11879
- Fix the torch version parsing logic by @houseroad in https://github.com/vllm-project/vllm/pull/15857
- [VLM] Remove
BaseProcessingInfo.get_mm_max_tokens_per_itemby @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16408 - [TPU][V1] Use
language_modelinterface for getting text backbone in MM by @NickLucche in https://github.com/vllm-project/vllm/pull/16410 - Improve configs -
ParallelConfigby @hmellor in https://github.com/vllm-project/vllm/pull/16332 - [V1] Set structured output backend to
autoby default by @russellb in https://github.com/vllm-project/vllm/pull/15724 - [V1][Spec Decode] Eagle Model loading by @LiuXiaoxuanPKU in https://github.com/vllm-project/vllm/pull/16035
- [Bugfix] Fix bug when dataset is json by @Chenyaaang in https://github.com/vllm-project/vllm/pull/15899
- [Model] Reduce redundant computations in mamba2 blocks for Bamba-9B by @cyang49 in https://github.com/vllm-project/vllm/pull/15423
- [V1] Zero-copy tensor/ndarray serialization/transmission by @njhill in https://github.com/vllm-project/vllm/pull/13790
- [VLM] Avoid unnecessary dummy multimodal data during processing by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16416
- [Bugfix] Fix output token length check logic by @eeslook in https://github.com/vllm-project/vllm/pull/16419
- [TPU][V1] Disable per-request seed/Generator by @NickLucche in https://github.com/vllm-project/vllm/pull/16172
- Fix range_ratio Bug in RandomDataset by @jadewang21 in https://github.com/vllm-project/vllm/pull/16126
- check input length of sonnet samples by @alexey-belyakov in https://github.com/vllm-project/vllm/pull/16423
- update benchmark_serving_structured_output to include auto backend by @Chenyaaang in https://github.com/vllm-project/vllm/pull/16438
- [Llama4] Enable attention temperature tuning by default for long context (>32k) by @sarckk in https://github.com/vllm-project/vllm/pull/16439
- Update supported_hardware.md for TPU INT8 by @mgoin in https://github.com/vllm-project/vllm/pull/16437
- [Bugfix][VLM] Fix failing Phi-4-MM multi-images tests and add vision-speech test by @Isotr0py in https://github.com/vllm-project/vllm/pull/16424
- [CPU][Bugfix] Fix CPU docker issues by @bigPYJ1151 in https://github.com/vllm-project/vllm/pull/16454
- [Bugfix] Don't set an upper bound on repetition penalty by @alex-jw-brooks in https://github.com/vllm-project/vllm/pull/16403
- Revert "[Model] use AutoWeightsLoader for deepseek_v2, internlm2" by @DefTruth in https://github.com/vllm-project/vllm/pull/16453
- [Core][LoRA][1/N] Add LoRA for EncoderDecoderModelRunner by @jeejeelee in https://github.com/vllm-project/vllm/pull/15990
- Enforce valid max_num_batched_tokens when disable_chunked_mm_input=True by @mgoin in https://github.com/vllm-project/vllm/pull/16447
- [Misc] Raise error for V1 not supporting Long LoRA. by @jeejeelee in https://github.com/vllm-project/vllm/pull/16415
- [Misc] update api_client example by @reidliu41 in https://github.com/vllm-project/vllm/pull/16459
- Don't install triton on
ppc64leplatform by @hmellor in https://github.com/vllm-project/vllm/pull/16470 - [Kernel] support merge_attn_states CUDA kernel, 3x speedup by @DefTruth in https://github.com/vllm-project/vllm/pull/16173
- [Bugfix] Fix bugs of running Quark quantized models by @cha557 in https://github.com/vllm-project/vllm/pull/16236
- [Hardware][Intel-Gaudi] Multi-step scheduling implementation for HPU by @tzielinski-habana in https://github.com/vllm-project/vllm/pull/12779
- Fix erroneous "model doesn't support compile" warning by @zou3519 in https://github.com/vllm-project/vllm/pull/16486
- [TPU][V1] Make
--disable_chunked_mm_inputmandatory for serving MM models by @NickLucche in https://github.com/vllm-project/vllm/pull/16483 - [Kernel] Support W8A8 channel-wise weights and per-token activations in triton fused_moe_kernel by @mgoin in https://github.com/vllm-project/vllm/pull/16366
- [Doc] Document InternVL3 support by @Isotr0py in https://github.com/vllm-project/vllm/pull/16495
- [Bugfix] handle alignment of encoder_seq_lens in mllama.py by @tjohnson31415 in https://github.com/vllm-project/vllm/pull/14784
- Improve configs -
LoadConfigby @hmellor in https://github.com/vllm-project/vllm/pull/16422 - [Frontend] Added chat templates for LLaMa4 pythonic tool calling by @yeqcharlotte in https://github.com/vllm-project/vllm/pull/16463
- [Kernel] Add tuned FusedMoE kernel config for Llama4 Scout, TP=8 on H100 by @sarckk in https://github.com/vllm-project/vllm/pull/16488
- Update openai_compatible_server.md by @Chr1st1anSears in https://github.com/vllm-project/vllm/pull/16507
- [Bugfix] clean up duplicated code by @lengrongfu in https://github.com/vllm-project/vllm/pull/16485
- Bugfix for PixtralHF models without spatial_merge_size by @mgoin in https://github.com/vllm-project/vllm/pull/16513
- [Doc] Fix link to vLLM blog by @terrytangyuan in https://github.com/vllm-project/vllm/pull/16519
- [CI][Bugfix] Add mistral_tool_use to Ci by @mgoin in https://github.com/vllm-project/vllm/pull/16517
- [BugFix] Handle non-contiguous tensors properly when serializing by @njhill in https://github.com/vllm-project/vllm/pull/16492
- [Doc] Update Llama4 Model Names in Supported Models by @yeqcharlotte in https://github.com/vllm-project/vllm/pull/16509
- Optimized topk for topk=1 (Llama-4) by @mgoin in https://github.com/vllm-project/vllm/pull/16512
- [Feature][V1] Add xgrammar to support minLength, maxLength with test by @leon-seidel in https://github.com/vllm-project/vllm/pull/16516
- [Frontend] support matryoshka representation / support embedding API dimensions by @noooop in https://github.com/vllm-project/vllm/pull/16331
- fix: spelling by @ezhoureal in https://github.com/vllm-project/vllm/pull/16466
- [Misc] Update chat utils tests by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/16520
- [Misc] Openai transcription client example use same Whisper model by @NickLucche in https://github.com/vllm-project/vllm/pull/16487
- [V1] Enable multi-input by default by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/15799
- [MISC] Make GroupCoordinator compatible with out-of-tree devices by @ji-huazhong in https://github.com/vllm-project/vllm/pull/16464
- [Misc] Delete redundant code by @jeejeelee in https://github.com/vllm-project/vllm/pull/16530
- Fix syntaxWarning: invalid escape sequence '\s' by @DamonFool in https://github.com/vllm-project/vllm/pull/16532
- [Perf] Optimize Preparing Inputs for GPU Model Runner by @SnowCharmQ in https://github.com/vllm-project/vllm/pull/16484
- [Bugfix] Validate logit biases to prevent out of vocab ids crashing engine by @rymc in https://github.com/vllm-project/vllm/pull/16529
- [V1][Spec Decode] KV cache slots for eagle heads by @LiuXiaoxuanPKU in https://github.com/vllm-project/vllm/pull/16370
- Enable PTPC FP8 for CompressedTensorsW8A8Fp8MoEMethod (triton fused_moe) by @mgoin in https://github.com/vllm-project/vllm/pull/16537
- [Benchmark][Bugfix] Fix SonnetDataset default values in benchmark_throughput.py by @JenZhao in https://github.com/vllm-project/vllm/pull/16556
- [Core][V0] Enable regex support with xgrammar by @russellb in https://github.com/vllm-project/vllm/pull/13228
Full Changelog: https://github.com/vllm-project/vllm/compare/v0.8.3...v0.8.4