llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2025-04-22 17:36:04 +00:00

Author	SHA1	Message	Date
Johannes Gäßler	f6711cef44	CUDA: determine FA parallel blocks at runtime	2025-03-16 14:36:57 +01:00
Lucas Moura Belo	3d652bfddf	readme : update bindings (#12229 )	2025-03-06 21:15:13 +02:00
Johannes Gäßler	5220a16d18	CUDA: fix FA logic for PTX 7.0 and CC >= 7.5 (#12222 )	2025-03-06 18:45:09 +01:00
David Huang	3ffbbd5ce1	HIP: rocWMMA documentation and enabling in workflow builds (#12179 ) * Enable rocWMMA for Windows CI build * Enable for Ubuntu * GGML_HIP_ROCWMMA_FATTN documentation work	2025-03-06 14:14:11 +01:00
Olivier Chafik	42994048a3	update function-calling.md w/ template override for functionary-small-v3.2 (#12214 )	2025-03-06 09:03:31 +00:00
Aaron Teo	e9b2f84f14	llava: add big-endian conversion for image encoder (#12218 ) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>	2025-03-06 09:33:21 +01:00
uvos	e721c05c93	HIP/CUDA: set the paramerter value in maintain_cuda_graph instead of replaceing it. (#12209 ) This avoids conflict with internal cuda/hip runtimes memory managment behavior. b4837	2025-03-06 08:20:52 +01:00
Han Yin	57b6abf85a	android : fix KV cache log message condition (#12212 ) b4836	2025-03-06 08:22:49 +02:00
Henry Linjamäki	94bb63e4f0	opencl : fix buffer alignment (#12197 ) Fix the following error: ``` ggml-alloc.c:99: not enough space in the buffer ggml_tallocr_alloc: not enough space in the buffer to allocate blk.17.ffn_down.weight (needed 27525120, available 27521024) ``` which occurs when `ggml_backend_opencl_context::alignment` is larger than `cl_ptr_base` (hard-coded to `0x1000`). Also, fix `ggml_backend_opencl_context::alignment` was set to `CL_DEVICE_MEM_BASE_ADDR_ALIGN` which was treated as bytes but the value is reported in bits. b4835	2025-03-06 02:33:40 +01:00
Henry Linjamäki	f79243992c	opencl : fix `ulong` kernel args were set from `int` variables (#12174 ) ... which left garbage bits in the upper half of the kernel args. This caused segmentation faults when running PoCL. b4834	2025-03-06 02:31:14 +01:00
simon886212	ed4ce0dda2	opencl : fix profile-related errors (#12095 ) Co-authored-by: ubuntu <ubuntu@localhost.localdomain> b4833	2025-03-06 02:30:05 +01:00
Rémy O	07d1572347	ggml-cpu: Faster IQ1 mul_mat_vec on AVX2 using BMI2 instructions (#12154 ) * ggml-cpu: Faster IQ1 mul_mat_vec on AVX2 using BMI2 instructions * cmake: Add GGML_BMI2 build option * ggml: enable BMI2 on relevant CPU variants * ggml-cpu: include BMI2 in backend score * ggml-cpu: register BMI2 in ggml_backend_cpu_get_features * ggml-cpu: add __BMI2__ define when using MSVC b4832	2025-03-06 02:26:10 +01:00
Akarshan Biswas	5e43f104cc	SYCL: Disable f16 Unary OPs as not supported by the kernels (#12201 ) b4831	2025-03-05 16:58:23 +01:00
Plamen Minev	16e4b22c5e	ggml : fix GGMLMetalClass ODR (#12200 ) -- it might happen if ggml is loaded from 2 separate libraries since each one of them will expose the class. This is more of a guard since we want to use only Metal as embedded library and don't care about the other case. b4830	2025-03-05 17:16:01 +02:00
Daniel Bevenius	074c4fd39d	ci : add fetch-depth to xcframework upload (#12195 ) This commit adds the fetch-depth: 0 option to the checkout action in the build.yml workflow file (0 meaning that it fetches the complete history). The default value is 1 when not specified which only fetches the latest commit. This is necessary to ensure that `git rev-list --count HEAD` counts the total number of commits in the history. Currently because the default is being used the name of the xcframework artifact is always llama-b1-xcframework. b4829	2025-03-05 14:16:40 +01:00
Olivier Chafik	669912d9a5	`tool-call`: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034 ) * sampler: turn lazy grammar trigger words to regexes * add scripts/tool_bench.sh & .py * constrain llama json output regardless of function name if matches at beginning * update relaxed newline space rule in grammar tests * support add_generation_prompt query parameter (useful for /apply_template) * Update src/llama-grammar.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-03-05 13:05:13 +00:00
Daniel Bevenius	fa31c438e0	ci : fix xcframework artifact tag (#12191 ) The commit add the name parameter to the upload-artifact action to ensure that the artifact is uploaded with the correct name. The motivation for this is that currently the uploaded xcframework is named as llama-b1-xcframework.zip. With this change the name of this artifact should contain the build number like the other artifacts. b4827	2025-03-05 10:22:29 +01:00
Daniel Bevenius	3ccbfe5a71	ci : remove xframework upload (#12190 ) * ci : remove xframework upload This commit removes the upload of the xframework zip file as an artifact. The motivation for this change is that the xframework zip file is currently being uploaded as part of strategy and will therefore be attempted to be uploaded multiple times and will fail the build. The uploading should be moved to somewhere else in the build to avoid this. * ci : add xcframework upload to macos-latest job b4826	2025-03-05 08:34:02 +01:00
Clauszy	06a92a193a	server : fix cache reuse logic (#12161 ) The first kv shift offsets the positions of all tokens after head_c. When using llama_kv_cache_seq_rm next, using head_c will remove the valid tokens because their positions have already been offset.	2025-03-05 09:25:45 +02:00
Daniel Bevenius	a057897ad4	llama : add xcframework build script (#11996 ) * llama : add xcframework build script This commit adds a script to build an XCFramework for Apple ios, macos, visionos, and tvos platforms. The generated XCFramework can then be added to a project and used in the same way as a regular framework. The llama.swiftui example project has been updated to use the XCFramework and can be started using the following command: ```console $ open examples/llama.swiftui/llama.swiftui.xcodeproj/ ``` Refs: https://github.com/ggml-org/llama.cpp/issues/10747 * examples : remove llama.cpp (source dir ref) from project.pbxproj This commit removes the reference to llama.cpp from the project.pbxproj file since Package.swift has been removed. * ci : updated build.yml to use build-xcframework.sh * ci : add xcframework build to github releases This commit adds the ability to create a GitHub release with the xcframework build artifact. * scripts : add apple app validation scripts This commit adds scripts that can validate the iOS, macOS, tvOS, and VisionOS applications. The scripts create a simple test app project, copy the llama.xcframework to the test project, build and archive the app, create an IPA from the archive, and validate the IPA using altool. The motivation for this is to provide some basic validation and hopefully avoid having to manually validate apps in Xcode. * llama : remove Package.swift This commit removes the Package.swift file, as we are now building an XCFramework for the project. * llama : remove Sources and spm-headers directories * llama : use TargetConditionals.h for visionOS/tvOS b4824	2025-03-05 06:30:31 +01:00
mgroeber9110	5bbe6a9fe9	ggml : portability fixes for VS 2017 (#12150 ) * Add include files for std::min/max and std::toupper/tolower * win32: move _USE_MATH_DEFINES before includes to ensure M_PI is defined * Use GGML_RESTRICT instead of "restrict" keyword everywhere, and use "__restrict" in MSVC plain C mode * win32: only use __restrict in MSVC if C11/C17 support is not enabled --------- Co-authored-by: Marcus Groeber <Marcus.Groeber@cerence.com> b4823	2025-03-04 18:53:26 +02:00
Georgi Gerganov	20a9b8f5e1	readme : fix roadmap link (#12185 )	2025-03-04 18:42:44 +02:00
Sigbjørn Skjæret	56d7a9f812	main: allow preloading conversation with -p and add -st / --single-turn (#12145 ) * Add chat template formatting to -no-cnv * only enable prompt formatting if explicitly enabled * add -st / --single-turn * add --single-turn and -p in conversation mode * fix -sys + -p * reword warning * small readability change and fix (long) outdated example usage * only activate single turn in conversation mode b4821	2025-03-04 12:19:39 -04:00
Olivier Chafik	1a24c4621f	`server`: fix deadly typo in response_format.json_schema.schema handling (#12168 ) b4820	2025-03-04 08:24:07 +02:00
David Huang	becade5de7	HIP: implement FlashAttention via rocWMMA for CDNA and RDNA3+ (#12032 ) Adds GGML_HIP_ROCWMMA_FATTN and rocwmma header check Adds rocWMMA support to fattn-wmma-f16 --- Signed-off-by: Carl Klemm <carl@uvos.xyz> Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Ben Jackson <ben@ben.com> b4819	2025-03-03 22:10:54 +01:00
Georgi Gerganov	dfd6b2c0be	sync : ggml ggml-ci b4818	2025-03-03 18:18:11 +02:00
cmdr2	b64d7cc272	cuda: unary ops as float + de-duplicate (ggml/1130)	2025-03-03 18:18:11 +02:00
Georgi Gerganov	3d1cf3cf33	sync : ggml ggml-ci	2025-03-03 18:18:11 +02:00
cmdr2	0cbee131ad	cuda/vulkan: specify fp32-only support for some operations in supports_op (ggml/1129) ggml-ci	2025-03-03 18:18:11 +02:00
Georgi Gerganov	8371d44595	sync : ggml ggml-ci	2025-03-03 18:18:11 +02:00
cmdr2	87abb7e903	cuda/cpu: Increase support for fp16 unary operations (ggml/1125) * Support fp16 unary operations in the CUDA backend * cpu: increase fp16 support for unary operators in the CPU backend * cuda: increase fp16 support for unary operators in the CUDA backend * Add test cases for fp16 unary operators * metal: update supports_op for unary operators that don't support fp16, to prevent test-backend-ops from failing * metal: fix PR comments for unary op support after fp16 unary tests	2025-03-03 18:18:11 +02:00
Diego Devesa	6d4c23b81b	whisper : support GGML_BACKEND_DL (whisper/2843) * whisper : support GGML_BACKEND_DL * fix DTW crash * whisper.objc : fix build - add ggml-cpp.h --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-03-03 18:18:11 +02:00
midnight	6512a90037	cmake : fix compile assumptions for power9/etc (whisper/2777) * Add small comment re: VSX to readme Co-authored-by: midnight <midnight@example.com>	2025-03-03 18:18:11 +02:00
petterreinholdtsen	4512055792	Told cmake to install ggml-cpp.h as a public header file. (ggml/1126) It is used by Whisper talk-llama example. Co-authored-by: Petter Reinholdtsen <pere@debian.org>	2025-03-03 18:18:11 +02:00
cmdr2	f54a4ba11e	Support pure float16 add/sub/mul/div operations in the CUDA (and CPU) backend (ggml/1121) * Support float16-to-float16 add/sub/mul/div operations in the CUDA backend * Add fp16 support for add/sub/mul/div on the CPU backend * Add test cases for fp16 add/sub/mul/div	2025-03-03 18:18:11 +02:00
Georgi Gerganov	aede2074f6	scripts : sync-ggml-am.sh fix	2025-03-03 18:18:11 +02:00
Daniel Bevenius	2679c3b55d	ci : set GITHUB_ACTION env var for server tests (#12162 ) This commit tries to address/improve an issue with the server tests which are failing with a timeout. Looking at the logs it seems like they are timing out after 12 seconds: ``` FAILED unit/test_chat_completion.py::test_completion_with_json_schema[False-json_schema0-6-"42"] - TimeoutError: Server did not start within 12 seconds ``` This is somewhat strange as in utils.py we have the following values: ```python DEFAULT_HTTP_TIMEOUT = 12 if "LLAMA_SANITIZE" in os.environ or "GITHUB_ACTION" in os.environ: DEFAULT_HTTP_TIMEOUT = 30 def start(self, timeout_seconds: int \| None = DEFAULT_HTTP_TIMEOUT) -> None: ``` It should be the case that a test running in a github action should have a timeout of 30 seconds. However, it seems like this is not the case. Inspecting the logs from the CI job we can see the following environment variables: ```console Run cd examples/server/tests 2 cd examples/server/tests 3 ./tests.sh 4 shell: /usr/bin/bash -e {0} 5 env: 6 LLAMA_LOG_COLORS: 1 7 LLAMA_LOG_PREFIX: 1 8 LLAMA_LOG_TIMESTAMPS: 1 9 LLAMA_LOG_VERBOSITY: 10 10 pythonLocation: /opt/hostedtoolcache/Python/3.11.11/x64 ``` This probably does not address the underlying issue that the servers that are providing the models to be downloaded occasionally take a longer time to response but might improve these situations in some cases.	2025-03-03 16:17:36 +01:00
dm4	c43af9276b	tts: add speaker file support (#12048 ) * tts: add speaker file support Signed-off-by: dm4 <sunrisedm4@gmail.com> * tts: handle outetts-0.3 * tts : add new line in error message --------- Signed-off-by: dm4 <sunrisedm4@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> b4806	2025-03-03 15:09:29 +02:00
Diego Devesa	d5c63cd7f9	test-backend-ops : add option -p to filter by op params (#12155 ) b4805	2025-03-03 14:00:46 +01:00
ag2s20150909	9660ffef58	ggml : fix kleidiai build (#12159 ) The libggml API has changed, but this has not been updated. b4804	2025-03-03 13:54:08 +01:00
Eric Curtin	c950a1f692	Adding UTF-8 support to llama.cpp (#12111 ) For emojis, non-alpha characters, etc. Signed-off-by: Eric Curtin <ecurtin@redhat.com> b4803	2025-03-03 12:44:56 +00:00
Xuan-Son Nguyen	7b69003af7	webui : add ?m=... and ?q=... params (#12148 ) * webui : add ?m=... and ?q=... params * also clear prefilledMessage variable * better approach * fix comment * test: bump timeout on GITHUB_ACTION	2025-03-03 11:42:45 +01:00
Akarshan Biswas	ece9745bb8	SYCL: Move CPY kernels to a separate file and add few missing kernels (#12133 ) * SYCL: refactor and move cpy kernels to a separate file * Add few missing cpy kernels * refactor and add debug logs b4801	2025-03-03 11:07:22 +01:00
Diego Devesa	cc473cac7c	ggml-backend : keep paths in native string type when possible (#12144 ) b4800	2025-03-02 22:11:00 +01:00
Sigbjørn Skjæret	14dec0c2f2	main: use jinja chat template system prompt by default (#12118 ) * Use jinja chat template system prompt by default * faster conditional order * remove nested ternary --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> b4799	2025-03-02 14:53:48 +01:00
Sigbjørn Skjæret	1782cdfed6	main: update outdated system prompt message (followup to #12131 ) (#12132 ) * Update outdated message * wording Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> b4798	2025-03-01 15:22:27 +01:00
Sigbjørn Skjæret	45a8e76745	common : add --system-prompt parameter, replace behavior of -p in conversation mode (#12131 ) * Add --system-prompt parameter * use user defined system prompt * clarify Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> * add warning * clarify Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> --------- Co-authored-by: Xuan-Son Nguyen <thichthat@gmail.com> b4797	2025-03-01 13:56:45 +01:00
Erik Scholz	80c41ddd8f	CUDA: compress mode option and default to size (#12029 ) cuda 12.8 added the option to specify stronger compression for binaries, so we now default to "size". b4796	2025-03-01 12:57:22 +01:00
Vivian	2cc4a5e44a	webui : minor typo fixes (#12116 ) * fix typos and improve menu text clarity * rename variable trimedValue to trimmedValue * add updated index.html.gz * rebuild --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co>	2025-03-01 11:15:09 +01:00
Xuan-Son Nguyen	06c2b1561d	convert : fix Norway problem when parsing YAML (#12114 ) * convert : fix Norway problem when parsing YAML * Update gguf-py/gguf/metadata.py * add newline at correct place	2025-02-28 17:44:46 +01:00

1 2 3 4 5 ...

4843 Commits