-
Notifications
You must be signed in to change notification settings - Fork 31
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
bug: Gemma 4 26B A4B temp-0 decode is non-deterministic run-to-run and corrupts near-tied tokens independently of the fused MoE kernel
priority:highHigh priorityHigh prioritystatus:in-progressCurrently being worked onCurrently being worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#910 In lablup/mlxcel;Epic: Fused paged decode, sorting-free sampling, and serving performance techniques
area:architectureArchitecture and code structure changesArchitecture and code structure changespriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#909 In lablup/mlxcel;Mixed prefill/decode step execution: design spike and prototype
area:architectureArchitecture and code structure changesArchitecture and code structure changespriority:lowLow priorityLow prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#908 In lablup/mlxcel;MLA matrix-absorbed decode path with compressed-latent KV cache for DeepSeek-family models
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#907 In lablup/mlxcel;Shape-bucketed kernel autotuner and cold-L2 benchmark methodology
area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#906 In lablup/mlxcel;Fused residual-add RMSNorm and fused RoPE + KV-append decode kernels
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#905 In lablup/mlxcel;Fused sparse-attention decode via page indirection for DSA and block-sparse models
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:mediumMedium priorityMedium prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:performancePerformance improvementsPerformance improvementsStatus: Open.#904 In lablup/mlxcel;Cascade attention: compute shared prompt prefixes once per decode batch
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:performancePerformance improvementsPerformance improvementsStatus: Open.#903 In lablup/mlxcel;Distribution-preserving speculative-decoding acceptance (chain speculative sampling)
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:mediumMedium priorityMedium prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#902 In lablup/mlxcel;Sorting-free top-k/top-p/min-p sampling with dual-pivot rejection kernels
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:performancePerformance improvementsPerformance improvementsStatus: Open.#901 In lablup/mlxcel;Softmax-free GPU sampling: Gumbel-max categorical sampling kernel
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#900 In lablup/mlxcel;Route production batched paged decode through the fused v2 kernel, retiring the gather-then-SDPA hot path
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:performancePerformance improvementsPerformance improvementsStatus: Open.#899 In lablup/mlxcel;