Skip to content
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching — Yilong Zhao (2024) | RDL Network