sync : ggml #3148

ggerganov · 2025-05-13T10:12:12Z

No description provided.

…/13343) * sycl: fixed non-contiguous src1 mul_mats (nc and batched) * Fixed wrong static_cast inside kernel

This assert fired running Qwen_Qwen3-30B-A3B-Q2_K.gguf: GGML_ASSERT(nei0 * nei1 <= 3072); The tensor is 8 x 512. Increase this array size to accommodate.

* rpc : add rpc_msg_set_tensor_hash_req Use a dedicated struct for the request of RPC_CMD_SET_TENSOR_HASH which makes the code cleaner. * fix

* CUDA: FA support for Deepseek (Ampere or newer) * do loop unrolling via C++ template

…858) * sycl : Implemented reorder Q4_0 mmvq Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> * sycl : Fixed mmvq being called when reorder is disabled * sycl : Improved comments in the quants header Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> * Use static_assert * safe_div -> ceil_div * Clarify qi comment * change the reorder tensor from init to execute OP * dbg * Undo changes to test-backend-ops * Refactor changes on top of q4_0 reorder fix * Missing Reverts * Refactored opt_for_reorder logic to simplify code path * Explicit inlining and unroll * Renamed mul_mat_algo enum for consistency --------- Signed-off-by: Alberto Cabrera <alberto.cabrera@codeplay.com> Co-authored-by: romain.biessy <romain.biessy@codeplay.com>

* vulkan: scalar flash attention implementation * vulkan: always use fp32 for scalar flash attention * vulkan: use vector loads in scalar flash attention shader * vulkan: remove PV matrix, helps with register usage * vulkan: reduce register usage in scalar FA, but perf may be slightly worse * vulkan: load each Q value once. optimize O reduction. more tuning * vulkan: support q4_0/q8_0 KV in scalar FA * CI: increase timeout to accommodate newly-supported tests * vulkan: for scalar FA, select between 1 and 8 rows * vulkan: avoid using Float16 capability in scalar FA

…ma4 400B (llama/13386)

* ggml-cpu: Integrate fp32=bf16xbf16 SME KleidiAI kernel Signed-off-by: Dan Johansson <dan.johansson@arm.com> * * code review fixes Signed-off-by: Dan Johansson <dan.johansson@arm.com> * * adds a comment that clarifies barrier usage Signed-off-by: Dan Johansson <dan.johansson@arm.com> --------- Signed-off-by: Dan Johansson <dan.johansson@arm.com> Co-authored-by: Charles Xu <charles.xu@arm.com>

* llama/ggml: add LLM training support more compact progress bar llama_save_model_to_file llama_opt_param_filter ggml_graph_dup force_grads refactor ggml_opt, fix test-opt * remove logits_all * refactor CUDA implementation for ACC * reset graph at beginning of opt period

ggml-ci

Alcpz and others added 21 commits May 13, 2025 13:05

sycl: addressing non-contiguous src1 mul_mats (nc and batched) (llama…

0c4a229

…/13343) * sycl: fixed non-contiguous src1 mul_mats (nc and batched) * Fixed wrong static_cast inside kernel

vulkan: Allow up to 4096 elements for mul_mat_id row_ids (llama/13326)

19d8d9a

This assert fired running Qwen_Qwen3-30B-A3B-Q2_K.gguf: GGML_ASSERT(nei0 * nei1 <= 3072); The tensor is 8 x 512. Increase this array size to accommodate.

rpc : add rpc_msg_set_tensor_hash_req (llama/13353)

00c8056

* rpc : add rpc_msg_set_tensor_hash_req Use a dedicated struct for the request of RPC_CMD_SET_TENSOR_HASH which makes the code cleaner. * fix

CUDA: fix crash on large batch size for MoE models (llama/13384)

f8c75dc

CUDA: FA support for Deepseek (Ampere or newer) (llama/13306)

aef59f4

* CUDA: FA support for Deepseek (Ampere or newer) * do loop unrolling via C++ template

CUDA: fix FlashAttention on Turing (llama/13415)

0444566

CUDA: fix race conditions FlashAttention kernels (llama/13438)

86dece9

Add --no-op-offload to improve -ot pp perf in MoE models like lla…

0b1962a

…ma4 400B (llama/13386)

CUDA: fix crash with partial offloading of MoE (llama/13439)

c426829

enable dpcpp nightly builds with libraries (llama/13406)

882d975

CUDA: fix misaligned synchronization in FA (llama/13469)

8264872

opencl: remove unnecessary assert for add (llama/13257)

43a59ec

metal : optimize MoE for large batches (llama/13388)

926e06d

ggml : add mrope kernel for metal (llama/13457)

79fb43e

sync : ggml

89970b9

ggml-ci

whisper : update to ggml-backend changes (#0)

6975380

ggml-ci

talk-llama : sync llama.cpp

bff8dc2

ggml-ci

danbev approved these changes May 13, 2025

View reviewed changes

ggerganov merged commit f890560 into master May 13, 2025
60 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

sync : ggml #3148

sync : ggml #3148

Uh oh!

ggerganov commented May 13, 2025

Uh oh!

Uh oh!

Uh oh!

sync : ggml #3148

sync : ggml #3148

Uh oh!

Conversation

ggerganov commented May 13, 2025

Uh oh!

Uh oh!

Uh oh!