0.6.4
Highlights
- Significant progress in V1 engine core refactor (#9826, #10135, #10288, #10211, #10225, #10228, #10268, #9954, #10272, #9971, #10224, #10166, #9289, #10058, #9888, #9972, #10059, #9945, #9679, #9871, #10227, #10245, #9629, #10097, #10203, #10148). You can checkout more details regarding the design and plan ahead in our recent meetup slides
- Signficant progress in
torch.compilesupport. Many models now support torch compile with TorchInductor. You can checkout our meetup slides for more details. (#9775, #9614, #9639, #9641, #9876, #9946, #9589, #9896, #9637, #9300, #9947, #9138, #9715, #9866, #9632, #9858, #9889)
Model Support
- New LLMs and VLMs: Idefics3 (#9767), H2OVL-Mississippi (#9747), Qwen2-Audio (#9248), Pixtral models in the HF Transformers format (#9036), FalconMamba (#9325), Florence-2 language backbone (#9555)
- New encoder-decoder embedding models: BERT (#9056), RoBERTa & XLM-RoBERTa (#9387)
- Expanded task support: Llama embeddings (#9806), Math-Shepherd (Mistral reward modeling) (#9697), Qwen2 classification (#9704), Qwen2 embeddings (#10184), VLM2Vec (Phi-3-Vision embeddings) (#9303), E5-V (LLaVA-NeXT embeddings) (#9576), Qwen2-VL embeddings (#9944)
- Add user-configurable
--taskparameter for models that support both generation and embedding (#9424) - Chat-based Embeddings API (#9759)
- Add user-configurable
- Tool calling parser for Granite 3.0 (#9027), Jamba (#9154), granite-20b-functioncalling (#8339)
- LoRA support for Granite 3.0 MoE (#9673), Idefics3 (#10281), Llama embeddings (#10071), Qwen (#9622), Qwen2-VL (#10022)
- BNB quantization support for Idefics3 (#10310), Mllama (#9720), Qwen2 (#9467, #9574), MiniCPMV (#9891)
- Unified multi-modal processor for VLM (#10040, #10044)
- Simplify model interface (#9933, #10237, #9938, #9958, #10007, #9978, #9983, #10205)
Hardware Support
- Gaudi: Add Intel Gaudi (HPU) inference backend (#6143)
- CPU: Add embedding models support for CPU backend (#10193)
- TPU: Correctly profile peak memory usage & Upgrade PyTorch XLA (#9438)
- Triton: Add Triton implementation for scaled_mm_triton to support fp8 and int8 SmoothQuant, symmetric case (#9857)
Performance
- Combine chunked prefill with speculative decoding (#9291)
fused_moePerformance Improvement (#9384)
Engine Core
- Override HF
config.jsonvia CLI (#5836) - Add goodput metric support (#9338)
- Move parallel sampling out from vllm core, paving way for V1 engine (#9302)
- Add stateless process group for easier integration with RLHF and disaggregated prefill (#10216, #10072)
Others
- Improvements to the pull request experience with DCO, mergify, stale bot, etc. (#9436, #9512, #9513, #9259, #10082, #10285, #9803)
- Dropped support for Python 3.8 (#10038, #8464)
- Basic Integration Test For TPU (#9968)
- Document the class hierarchy in vLLM (#10240), explain the integration with Hugging Face (#10173).
- Benchmark throughput now supports image input (#9851)
What's Changed
- [TPU] Fix TPU SMEM OOM by Pallas paged attention kernel by @WoosukKwon in https://github.com/vllm-project/vllm/pull/9350
- [Frontend] merge beam search implementations by @LunrEclipse in https://github.com/vllm-project/vllm/pull/9296
- [Model] Make llama3.2 support multiple and interleaved images by @xiangxu-google in https://github.com/vllm-project/vllm/pull/9095
- [Bugfix] Clean up some cruft in mamba.py by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/9343
- [Frontend] Clarify model_type error messages by @stevegrubb in https://github.com/vllm-project/vllm/pull/9345
- [Doc] Fix code formatting in spec_decode.rst by @mgoin in https://github.com/vllm-project/vllm/pull/9348
- [Bugfix] Update InternVL input mapper to support image embeds by @hhzhang16 in https://github.com/vllm-project/vllm/pull/9351
- [BugFix] Fix chat API continuous usage stats by @njhill in https://github.com/vllm-project/vllm/pull/9357
- pass ignore_eos parameter to all benchmark_serving calls by @gracehonv in https://github.com/vllm-project/vllm/pull/9349
- [Misc] Directly use compressed-tensors for checkpoint definitions by @mgoin in https://github.com/vllm-project/vllm/pull/8909
- [Bugfix] Fix vLLM UsageInfo and logprobs None AssertionError with empty token_ids by @CatherineSue in https://github.com/vllm-project/vllm/pull/9034
- [Bugfix][CI/Build] Fix CUDA 11.8 Build by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/9386
- [Bugfix] Molmo text-only input bug fix by @mrsalehi in https://github.com/vllm-project/vllm/pull/9397
- [Misc] Standardize RoPE handling for Qwen2-VL by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9250
- [Model] VLM2Vec, the first multimodal embedding model in vLLM by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9303
- [CI/Build] Test VLM embeddings by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9406
- [Core] Rename input data types by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/8688
- [Misc] Consolidate example usage of OpenAI client for multimodal models by @ywang96 in https://github.com/vllm-project/vllm/pull/9412
- [Model] Support SDPA attention for Molmo vision backbone by @Isotr0py in https://github.com/vllm-project/vllm/pull/9410
- Support mistral interleaved attn by @patrickvonplaten in https://github.com/vllm-project/vllm/pull/9414
- [Kernel][Model] Improve continuous batching for Jamba and Mamba by @mzusman in https://github.com/vllm-project/vllm/pull/9189
- [Model][Bugfix] Add FATReLU activation and support for openbmb/MiniCPM-S-1B-sft by @streaver91 in https://github.com/vllm-project/vllm/pull/9396
- [Performance][Spec Decode] Optimize ngram lookup performance by @LiuXiaoxuanPKU in https://github.com/vllm-project/vllm/pull/9333
- [CI/Build] mypy: Resolve some errors from checking vllm/engine by @russellb in https://github.com/vllm-project/vllm/pull/9267
- [Bugfix][Kernel] Prevent integer overflow in fp8 dynamic per-token quantize kernel by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/9425
- [BugFix] [Kernel] Fix GPU SEGV occurring in int8 kernels by @rasmith in https://github.com/vllm-project/vllm/pull/9391
- Add notes on the use of Slack by @terrytangyuan in https://github.com/vllm-project/vllm/pull/9442
- [Kernel] Add Exllama as a backend for compressed-tensors by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/9395
- [Misc] Print stack trace using
logger.exceptionby @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9461 - [misc] CUDA Time Layerwise Profiler by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/8337
- [Bugfix] Allow prefill of assistant response when using
mistral_commonby @sasha0552 in https://github.com/vllm-project/vllm/pull/9446 - [TPU] Call torch._sync(param) during weight loading by @WoosukKwon in https://github.com/vllm-project/vllm/pull/9437
- [Hardware][CPU] compressed-tensor INT8 W8A8 AZP support by @bigPYJ1151 in https://github.com/vllm-project/vllm/pull/9344
- [Core] Deprecating block manager v1 and make block manager v2 default by @KuntaiDu in https://github.com/vllm-project/vllm/pull/8704
- [CI/Build] remove .github from .dockerignore, add dirty repo check by @dtrifiro in https://github.com/vllm-project/vllm/pull/9375
- [Misc] Remove commit id file by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9470
- [torch.compile] Fine-grained CustomOp enabling mechanism by @ProExpertProg in https://github.com/vllm-project/vllm/pull/9300
- [Bugfix] Fix support for dimension like integers and ScalarType by @bnellnm in https://github.com/vllm-project/vllm/pull/9299
- [Bugfix] Add random_seed to sample_hf_requests in benchmark_serving script by @wukaixingxp in https://github.com/vllm-project/vllm/pull/9013
- [Bugfix] Print warnings related to
mistral_commontokenizer only once by @sasha0552 in https://github.com/vllm-project/vllm/pull/9468 - [Hardwware][Neuron] Simplify model load for transformers-neuronx library by @sssrijan-amazon in https://github.com/vllm-project/vllm/pull/9380
- Support
BERTModel(firstencoder-onlyembedding model) by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9056 - [BugFix] Stop silent failures on compressed-tensors parsing by @dsikka in https://github.com/vllm-project/vllm/pull/9381
- [Bugfix][Core] Use torch.cuda.memory_stats() to profile peak memory usage by @joerunde in https://github.com/vllm-project/vllm/pull/9352
- [Qwen2.5] Support bnb quant for Qwen2.5 by @blueyo0 in https://github.com/vllm-project/vllm/pull/9467
- [CI/Build] Use commit hash references for github actions by @russellb in https://github.com/vllm-project/vllm/pull/9430
- [BugFix] Typing fixes to RequestOutput.prompt and beam search by @njhill in https://github.com/vllm-project/vllm/pull/9473
- [Frontend][Feature] Add jamba tool parser by @tomeras91 in https://github.com/vllm-project/vllm/pull/9154
- [BugFix] Fix and simplify completion API usage streaming by @njhill in https://github.com/vllm-project/vllm/pull/9475
- [CI/Build] Fix lint errors in mistral tokenizer by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9504
- [Bugfix] Fix offline_inference_with_prefix.py by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/9505
- [Misc] benchmark: Add option to set max concurrency by @russellb in https://github.com/vllm-project/vllm/pull/9390
- [Model] Add user-configurable task for models that support both generation and embedding by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9424
- [CI/Build] Add error matching config for mypy by @russellb in https://github.com/vllm-project/vllm/pull/9512
- [Model] Support Pixtral models in the HF Transformers format by @mgoin in https://github.com/vllm-project/vllm/pull/9036
- [MISC] Add lora requests to metrics by @coolkp in https://github.com/vllm-project/vllm/pull/9477
- [MISC] Consolidate cleanup() and refactor offline_inference_with_prefix.py by @comaniac in https://github.com/vllm-project/vllm/pull/9510
- [Kernel] Add env variable to force flashinfer backend to enable tensor cores by @tdoublep in https://github.com/vllm-project/vllm/pull/9497
- [Bugfix] Fix offline mode when using
mistral_commonby @sasha0552 in https://github.com/vllm-project/vllm/pull/9457 - :bug: fix torch memory profiling by @joerunde in https://github.com/vllm-project/vllm/pull/9516
- [Frontend] Avoid creating guided decoding LogitsProcessor unnecessarily by @njhill in https://github.com/vllm-project/vllm/pull/9521
- [Doc] update gpu-memory-utilization flag docs by @joerunde in https://github.com/vllm-project/vllm/pull/9507
- [CI/Build] Add error matching for ruff output by @russellb in https://github.com/vllm-project/vllm/pull/9513
- [CI/Build] Configure matcher for actionlint workflow by @russellb in https://github.com/vllm-project/vllm/pull/9511
- [Frontend] Support simpler image input format by @yue-anyscale in https://github.com/vllm-project/vllm/pull/9478
- [Bugfix] Fix missing task for speculative decoding by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9524
- [Model][Pixtral] Optimizations for input_processor_for_pixtral_hf by @mgoin in https://github.com/vllm-project/vllm/pull/9514
- [Bugfix] Pass json-schema to GuidedDecodingParams and make test stronger by @heheda12345 in https://github.com/vllm-project/vllm/pull/9530
- [Model][Pixtral] Use memory_efficient_attention for PixtralHFVision by @mgoin in https://github.com/vllm-project/vllm/pull/9520
- [Kernel] Support sliding window in flash attention backend by @heheda12345 in https://github.com/vllm-project/vllm/pull/9403
- [Frontend][Misc] Goodput metric support by @Imss27 in https://github.com/vllm-project/vllm/pull/9338
- [CI/Build] Split up decoder-only LM tests by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9488
- [Doc] Consistent naming of attention backends by @tdoublep in https://github.com/vllm-project/vllm/pull/9498
- [Model] FalconMamba Support by @dhiaEddineRhaiem in https://github.com/vllm-project/vllm/pull/9325
- [Bugfix][Misc]: fix graph capture for decoder by @yudian0504 in https://github.com/vllm-project/vllm/pull/9549
- [BugFix] Use correct python3 binary in Docker.ppc64le entrypoint by @varad-ahirwadkar in https://github.com/vllm-project/vllm/pull/9492
- [Model][Bugfix] Fix batching with multi-image in PixtralHF by @mgoin in https://github.com/vllm-project/vllm/pull/9518
- [Frontend] Reduce frequency of client cancellation checking by @njhill in https://github.com/vllm-project/vllm/pull/7959
- [doc] fix format by @youkaichao in https://github.com/vllm-project/vllm/pull/9562
- [BugFix] Update draft model TP size check to allow matching target TP size by @njhill in https://github.com/vllm-project/vllm/pull/9394
- [Frontend] Don't log duplicate error stacktrace for every request in the batch by @wallashss in https://github.com/vllm-project/vllm/pull/9023
- [CI] Make format checker error message more user-friendly by using emoji by @KuntaiDu in https://github.com/vllm-project/vllm/pull/9564
- :bug: Fixup more test failures from memory profiling by @joerunde in https://github.com/vllm-project/vllm/pull/9563
- [core] move parallel sampling out from vllm core by @youkaichao in https://github.com/vllm-project/vllm/pull/9302
- [Bugfix]: serialize config instances by value when using --trust-remote-code by @tjohnson31415 in https://github.com/vllm-project/vllm/pull/6751
- [CI/Build] Remove unnecessary
fork_new_processby @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9484 - [Bugfix][OpenVINO] fix_dockerfile_openvino by @ngrozae in https://github.com/vllm-project/vllm/pull/9552
- [Bugfix]: phi.py get rope_theta from config file by @Falko1 in https://github.com/vllm-project/vllm/pull/9503
- [CI/Build] Replaced some models on tests for smaller ones by @wallashss in https://github.com/vllm-project/vllm/pull/9570
- [Core] Remove evictor_v1 by @KuntaiDu in https://github.com/vllm-project/vllm/pull/9572
- [Doc] Use shell code-blocks and fix section headers by @rafvasq in https://github.com/vllm-project/vllm/pull/9508
- support TP in qwen2 bnb by @chenqianfzh in https://github.com/vllm-project/vllm/pull/9574
- [Hardware][CPU] using current_platform.is_cpu by @wangshuai09 in https://github.com/vllm-project/vllm/pull/9536
- [V1] Implement vLLM V1 [1/N] by @WoosukKwon in https://github.com/vllm-project/vllm/pull/9289
- [CI/Build][LoRA] Temporarily fix long context failure issue by @jeejeelee in https://github.com/vllm-project/vllm/pull/9579
- [Neuron] [Bugfix] Fix neuron startup by @xendo in https://github.com/vllm-project/vllm/pull/9374
- [Model][VLM] Initialize support for Mono-InternVL model by @Isotr0py in https://github.com/vllm-project/vllm/pull/9528
- [Bugfix] Eagle: change config name for fc bias by @gopalsarda in https://github.com/vllm-project/vllm/pull/9580
- [Hardware][Intel CPU][DOC] Update docs for CPU backend by @zhouyuan in https://github.com/vllm-project/vllm/pull/6212
- [Frontend] Support custom request_id from request by @guoyuhong in https://github.com/vllm-project/vllm/pull/9550
- [BugFix] Prevent exporting duplicate OpenTelemetry spans by @ronensc in https://github.com/vllm-project/vllm/pull/9017
- [torch.compile] auto infer dynamic_arg_dims from type annotation by @youkaichao in https://github.com/vllm-project/vllm/pull/9589
- [Bugfix] fix detokenizer shallow copy by @aurickq in https://github.com/vllm-project/vllm/pull/5919
- [Misc] Make benchmarks use EngineArgs by @JArnoldAMD in https://github.com/vllm-project/vllm/pull/9529
- [Bugfix] Fix spurious "No compiled cutlass_scaled_mm ..." for W8A8 on Turing by @LucasWilkinson in https://github.com/vllm-project/vllm/pull/9487
- [BugFix] Fix metrics error for --num-scheduler-steps > 1 by @yuleil in https://github.com/vllm-project/vllm/pull/8234
- [Doc]: Update tensorizer docs to include vllm[tensorizer] by @sethkimmel3 in https://github.com/vllm-project/vllm/pull/7889
- [Bugfix] Generate exactly input_len tokens in benchmark_throughput by @heheda12345 in https://github.com/vllm-project/vllm/pull/9592
- [Misc] Add an env var VLLM_LOGGING_PREFIX, if set, it will be prepend to all logging messages by @sfc-gh-zhwang in https://github.com/vllm-project/vllm/pull/9590
- [Model] Support E5-V by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9576
- [Build] Fix
FetchContentmultiple build issue by @ProExpertProg in https://github.com/vllm-project/vllm/pull/9596 - [Hardware][XPU] using current_platform.is_xpu by @MengqingCao in https://github.com/vllm-project/vllm/pull/9605
- [Model] Initialize Florence-2 language backbone support by @Isotr0py in https://github.com/vllm-project/vllm/pull/9555
- [VLM] Post-layernorm override and quant config in vision encoder by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9217
- [Model] Add min_pixels / max_pixels to Qwen2VL as mm_processor_kwargs by @alex-jw-brooks in https://github.com/vllm-project/vllm/pull/9612
- [Bugfix] Fix
_init_vision_modelin NVLM_D model by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9611 - [misc] comment to avoid future confusion about baichuan by @youkaichao in https://github.com/vllm-project/vllm/pull/9620
- [Bugfix] Fix divide by zero when serving Mamba models by @tlrmchlsmth in https://github.com/vllm-project/vllm/pull/9617
- [Misc] Separate total and output tokens in benchmark_throughput.py by @mgoin in https://github.com/vllm-project/vllm/pull/8914
- [torch.compile] Adding torch compile annotations to some models by @CRZbulabula in https://github.com/vllm-project/vllm/pull/9614
- [Frontend] Enable Online Multi-image Support for MLlama by @alex-jw-brooks in https://github.com/vllm-project/vllm/pull/9393
- [Model] Add Qwen2-Audio model support by @faychu in https://github.com/vllm-project/vllm/pull/9248
- [CI/Build] Add bot to close stale issues and PRs by @russellb in https://github.com/vllm-project/vllm/pull/9436
- [Bugfix][Model] Fix Mllama SDPA illegal memory access for batched multi-image by @mgoin in https://github.com/vllm-project/vllm/pull/9626
- [Bugfix] Use "vision_model" prefix for MllamaVisionModel by @mgoin in https://github.com/vllm-project/vllm/pull/9628
- [Bugfix]: Make chat content text allow type content by @vrdn-23 in https://github.com/vllm-project/vllm/pull/9358
- [XPU] avoid triton import for xpu by @yma11 in https://github.com/vllm-project/vllm/pull/9440
- [Bugfix] Fix PP for ChatGLM and Molmo, and weight loading for Qwen2.5-Math-RM by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9422
- [V1][Bugfix] Clean up requests when aborted by @WoosukKwon in https://github.com/vllm-project/vllm/pull/9629
- [core] simplify seq group code by @youkaichao in https://github.com/vllm-project/vllm/pull/9569
- [torch.compile] Adding torch compile annotations to some models by @CRZbulabula in https://github.com/vllm-project/vllm/pull/9639
- [Kernel] add kernel for FATReLU by @jeejeelee in https://github.com/vllm-project/vllm/pull/9610
- [torch.compile] expanding support and fix allgather compilation by @CRZbulabula in https://github.com/vllm-project/vllm/pull/9637
- [Doc] Move additional tips/notes to the top by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9647
- [Bugfix]Disable the post_norm layer of the vision encoder for LLaVA models by @litianjian in https://github.com/vllm-project/vllm/pull/9653
- Increase operation per run limit for "Close inactive issues and PRs" workflow by @hmellor in https://github.com/vllm-project/vllm/pull/9661
- [torch.compile] Adding torch compile annotations to some models by @CRZbulabula in https://github.com/vllm-project/vllm/pull/9641
- [CI/Build] Fix VLM test failures when using transformers v4.46 by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9666
- [Model] Compute Llava Next Max Tokens / Dummy Data From Gridpoints by @alex-jw-brooks in https://github.com/vllm-project/vllm/pull/9650
- [Log][Bugfix] Fix default value check for
image_url.detailby @mgoin in https://github.com/vllm-project/vllm/pull/9663 - [Performance][Kernel] Fused_moe Performance Improvement by @charlifu in https://github.com/vllm-project/vllm/pull/9384
- [Bugfix] Remove xformers requirement for Pixtral by @mgoin in https://github.com/vllm-project/vllm/pull/9597
- [ci/Build] Skip Chameleon for transformers 4.46.0 on broadcast test #9675 by @khluu in https://github.com/vllm-project/vllm/pull/9676
- [Model] add a lora module for granite 3.0 MoE models by @willmj in https://github.com/vllm-project/vllm/pull/9673
- [V1] Support sliding window attention by @WoosukKwon in https://github.com/vllm-project/vllm/pull/9679
- [Bugfix] Fix compressed_tensors_moe bad config.strategy by @mgoin in https://github.com/vllm-project/vllm/pull/9677
- [Doc] Improve quickstart documentation by @rafvasq in https://github.com/vllm-project/vllm/pull/9256
- [Bugfix] Fix crash with llama 3.2 vision models and guided decoding by @tjohnson31415 in https://github.com/vllm-project/vllm/pull/9631
- [Bugfix] Steaming continuous_usage_stats default to False by @samos123 in https://github.com/vllm-project/vllm/pull/9709
- [Hardware][openvino] is_openvino --> current_platform.is_openvino by @MengqingCao in https://github.com/vllm-project/vllm/pull/9716
- Fix: MI100 Support By Bypassing Custom Paged Attention by @MErkinSag in https://github.com/vllm-project/vllm/pull/9560
- [Frontend] Bad words sampling parameter by @Alvant in https://github.com/vllm-project/vllm/pull/9717
- [Model] Add classification Task with Qwen2ForSequenceClassification by @kakao-kevin-us in https://github.com/vllm-project/vllm/pull/9704
- [Misc] SpecDecodeWorker supports profiling by @Abatom in https://github.com/vllm-project/vllm/pull/9719
- [core] cudagraph output with tensor weak reference by @youkaichao in https://github.com/vllm-project/vllm/pull/9724
- [Misc] Upgrade to pytorch 2.5 by @bnellnm in https://github.com/vllm-project/vllm/pull/9588
- Fix cache management in "Close inactive issues and PRs" actions workflow by @hmellor in https://github.com/vllm-project/vllm/pull/9734
- [Bugfix] Fix load config when using bools by @madt2709 in https://github.com/vllm-project/vllm/pull/9533
- [Hardware][ROCM] using current_platform.is_rocm by @wangshuai09 in https://github.com/vllm-project/vllm/pull/9642
- [torch.compile] support moe models by @youkaichao in https://github.com/vllm-project/vllm/pull/9632
- Fix beam search eos by @robertgshaw2-neuralmagic in https://github.com/vllm-project/vllm/pull/9627
- [Bugfix] Fix ray instance detect issue by @yma11 in https://github.com/vllm-project/vllm/pull/9439
- [CI/Build] Adopt Mergify for auto-labeling PRs by @russellb in https://github.com/vllm-project/vllm/pull/9259
- [Model][VLM] Add multi-video support for LLaVA-Onevision by @litianjian in https://github.com/vllm-project/vllm/pull/8905
- [torch.compile] Adding "torch compile" annotations to some models by @CRZbulabula in https://github.com/vllm-project/vllm/pull/9758
- [misc] avoid circular import by @youkaichao in https://github.com/vllm-project/vllm/pull/9765
- [torch.compile] add deepseek v2 compile by @youkaichao in https://github.com/vllm-project/vllm/pull/9775
- [Doc] fix third-party model example by @russellb in https://github.com/vllm-project/vllm/pull/9771
- [Model][LoRA]LoRA support added for Qwen by @jeejeelee in https://github.com/vllm-project/vllm/pull/9622
- [Doc] Specify async engine args in docs by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9726
- [Bugfix] Use temporary directory in registry by @DarkLight1337 in https://github.com/vllm-project/vllm/pull/9721
- [Frontend] re-enable multi-modality input in the new beam search implementation by @FerdinandZhong in https://github.com/vllm-project/vllm/pull/9427
- [Model] Add BNB quantization support for Mllama by @Isotr0py in https://github.com/vllm-project/vllm/pull/9720
- [Hardware] using current_platform.seed_everything by @wangshuai09 in https://github.com/vllm-project/vllm/pull/9785
- [Misc] Add metrics for request queue time, forward time, and execute time by @Abatom in https://github.com/vllm-project/vllm/pull/9659
- Fix the log to correct guide user to install modelscope by @tastelikefeet in https://github.com/vllm-project/vllm/pull/9793
- [Bugfix] Use host argument to bind to interface by @svenseeberg in https://github.com/vllm-project/vllm/pull/9798
- [Misc]: Typo fix: Renaming classes (casualLM -> causalLM) by @yannicks1 in https://github.com/vllm-project/vllm/pull/9801
- [Model] Add LlamaEmbeddingModel as an embedding Implementation of LlamaModel by @jsato8094 in https://github.com/vllm-project/vllm/pull/9806
- [CI][Bugfix] Skip chameleon for transformers 4.46.1 by @mgoin in https://github.com/vllm-project/vllm/pull/9808
- [CI/Build] mergify: fix rules for ci/build label by @russellb in https://github.com/vllm-project/vllm/pull/9804
- [MISC] Set label value to timestamp over 0, to keep track of recent history by @coolkp in https://github.com/vllm-project/vllm/pull/9777
- [Bugfix][Frontend] Guard against bad token ids by @joerunde in https://github.com/vllm-project/vllm/pull/9634
- [Model] tool calling support for ibm-granite/granite-20b-functioncalling by @wseaton in https://github.com/vllm-project/vllm/pull/8339
- [Docs] Add notes about Snowflake Meetup by @simon-mo in https://github.com/vllm-project/vllm/pull/9814
- [Bugfix] Fix prefix strings for quantized VLMs by @mgoin in https://github.com/vllm-project/vllm/pull/9772
- [core][distributed] fix custom allreduce in pytorch 2.5 by @youkaichao in https://github.com/vllm-project/vllm/pull/9815
- Update README.md by @LiuXiaoxuanPKU in https://github.com/vllm-project/vllm/pull/9819
- [Bugfix][VLM] Make apply_fp8_linear work with >2D input by @mgoin in https://github.com/vllm-project/vllm/pull/9812
- [ci/build] Pin CI dependencies version with pip-compile by @khluu in https://github.com/vllm-project/vllm/pull/9810
- [Bugfix] Fix multi nodes TP+PP for XPU by @yma11 in https://github.com/vllm-project/vllm/pull/8884
- [Doc] Add the DCO to CONTRIBUTING.md by @russellb in https://github.com/vllm-project/vllm/pull/9803
- [torch.compile] rework compile control with piecewise cudagraph by @youkaichao in
These notes run past the length kept in the archive. The rest is on the publisher’s page.