# Furiosa LLM

**URL:** https://forums.furiosa.ai/c/furiosa-llm/10.md

[Latest](https://forums.furiosa.ai/latest.md) · [Categories](https://forums.furiosa.ai/categories.md) · [Tags](https://forums.furiosa.ai/tags.md)

---

## [About the Furiosa LLM category](https://forums.furiosa.ai/t/about-the-furiosa-llm-category/222)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0

</div>

---

## [Furiosa-LLM 2026.3.0 has been released today!](https://forums.furiosa.ai/t/furiosa-llm-2026-3-0-has-been-released-today/418)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0\
**Last updated:** [June 30, 2026, 1:35pm UTC](https://forums.furiosa.ai/t/furiosa-llm-2026-3-0-has-been-released-today/418 "2026-06-30T13:35:42Z")

</div>

Hi folks, Furiosa-LLM 2026.3.0 has been released today, and this release includes a lot of updates and new model support. Please see Furiosa SDK 2026.3.0 release to learn more about the release. You can also find new m…

---

## [Throughput 성능 최적화를 위해 TP=1일 때, furiosa artifact를 만드는 과정에서 에러가 나왔습니다](https://forums.furiosa.ai/t/throughput-tp-1-furiosa-artifact/378)

<div class="topic-metadata">

**Author:** [@cake](https://forums.furiosa.ai/u/cake)\
**Replies:** 2\
**Last updated:** [December 5, 2025, 4:44am UTC](https://forums.furiosa.ai/t/throughput-tp-1-furiosa-artifact/378 "2025-12-05T04:44:48Z")

</div>

안녕하세요. 퓨리오사 레니게이드에서 빌드 중에 궁금한 것이 있어 문의드립니다. Throughput 성능 최적화 방법 문의 글을 참고하여 모델 artifact를 추가해 빌드를 시도했습니다. from furiosa\_llm.artifact.builder import ArtifactBuilder RELEASE\_PREFILL\_BUCKETS = \[ (1, 256), (1, 320), (1,…

---

## [Gemma 3 27b 모델 양자화 오류](https://forums.furiosa.ai/t/gemma-3-27b/368)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 4\
**Last updated:** [December 5, 2025, 4:22am UTC](https://forums.furiosa.ai/t/gemma-3-27b/368 "2025-12-05T04:22:39Z")

</div>

안녕하세요? 항상 빠른 대응에 감사드립니다. 현재 상태는 다음과 같습니다. llama 3 8b instruct 모델은 양자화가 잘 되고 있습니다. gemma 모델 양자화는 오류가 발생하고 있습니다. gemma 3 27b 모델을 양자화하려고 하니 transformers 버전이 낮다는 오류가 발생하였습니다. pip install --upgrade transformers 를 수행시…

---

## [Llama3.1-8B 모델 컴파일 관련 문의](https://forums.furiosa.ai/t/llama3-1-8b/376)

<div class="topic-metadata">

**Author:** [@gspark](https://forums.furiosa.ai/u/gspark)\
**Replies:** 2\
**Last updated:** [November 14, 2025, 5:39am UTC](https://forums.furiosa.ai/t/llama3-1-8b/376 "2025-11-14T05:39:19Z")

</div>

:one:컴파일 속도 관련 현재 컴파일 속도가 매우 느립니다. 특히 RELEASE\_PREFILL\_BUCKETS, RELEASE\_DECODE\_BUCKETS 설정을 아래와 같이 지정했을 때 컴파일 시간이 크게 증가했습니다. 해당 과정이 어떠한 과정을 수행하는 지에 대한 설명과 컴파일 속도를 개선할 수 있는 방법이나 권장 설정이 있을까요? RELEASE\_PREFILL\_BUCKETS = \[ …

---

## [Rngd 에서 추론 중 오류](https://forums.furiosa.ai/t/rngd/380)

<div class="topic-metadata">

**Author:** [@noa](https://forums.furiosa.ai/u/noa)\
**Replies:** 1\
**Last updated:** [November 14, 2025, 5:27am UTC](https://forums.furiosa.ai/t/rngd/380 "2025-11-14T05:27:41Z")

</div>

안녕하세요. RNGD 환경에서 LLM 추론을 돌리다 보면 중간에 진행이 멈추고, 에러가 뜨는 현상이 지속적으로 일어납니다. 모델 크기마다 데이터셋 토큰 길이마다 달랐던 것으로 보여졌습니다. 어떤경우 117번째 시도 라거나 특정 구간에서 멈추면 같은 옵션 및 모델로 다시 시도해도 117번째에서 에러가 뜨는 모습을 보였습니다. 현재 앨리스 클라우드 상에서 RNGD 1장 장착 되어있는 서버에…

---

## [Furiosa NPU 4개 중 1개가 사라지는 현상 관련 문의](https://forums.furiosa.ai/t/furiosa-npu-4-1/381)

<div class="topic-metadata">

**Author:** [@gspark](https://forums.furiosa.ai/u/gspark)\
**Replies:** 2\
**Last updated:** [November 13, 2025, 7:31am UTC](https://forums.furiosa.ai/t/furiosa-npu-4-1/381 "2025-11-13T07:31:17Z")

</div>

안녕하세요. ETRI 고성능컴퓨팅시스템연구실 박경서입니다. 현재 Furiosa NPU 4개를 사용 중인데, 실험을 수행하다가 아래와 같이 NPU 한 개가 시스템에서 인식되지 않는 문제가 발생했습니다. 이 상태에서 NPU 기반 추론 작업을 수행하면 오류가 발생합니다. 이와 같은 상황에서 어떤 방식으로 문제를 진단하고 해결해야 하는지 안내 부탁드립니다. 현재 Container를 대여하여…

---

## [Models with sliding window attention](https://forums.furiosa.ai/t/models-with-sliding-window-attention/372)

<div class="topic-metadata">

**Author:** [@huijong.jeong](https://forums.furiosa.ai/u/huijong.jeong)\
**Replies:** 3\
**Last updated:** [November 4, 2025, 2:50am UTC](https://forums.furiosa.ai/t/models-with-sliding-window-attention/372 "2025-11-04T02:50:08Z")

</div>

\[EN\] I want to benchmark LLM models involving sliding window attention via furiosa-llm. Unfortunately, my model’s architecture is not on your supported models list and and it seems like the only way to meet my goal is …

---

## [아티팩트 생성 시 오류 발생 이슈](https://forums.furiosa.ai/t/topic/370)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 10\
**Last updated:** [November 3, 2025, 2:16am UTC](https://forums.furiosa.ai/t/topic/370 "2025-11-03T02:16:48Z")

</div>

빠른 지원에 감사드리고 있습니다. 이번에는 다음과 같이 아티팩트 생성과 관련된 내용입니다. 아래와 같은 명령을 이용하여 llama 3.1 8b를 미세조정한 모델을 양자화한 후, 아티팩트를 만들고자 하였습니다. furiosa-llm build ./quantized\_llam318\_komedic ./komedic\_qt8b -tp 8 --max-seq-len-to-c…

---

## [Librenegade.so 오류](https://forums.furiosa.ai/t/librenegade-so/365)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 8\
**Last updated:** [October 20, 2025, 1:20pm UTC](https://forums.furiosa.ai/t/librenegade-so/365 "2025-10-20T13:20:23Z")

</div>

이미 이전에 llama 3.1 8B 모델로 양자화도 하고 아티팩트도 만들어서 실행도 확인했었는데요. 해당 모듈을 다시 실행시키니 librenegade.so를 찾을 수 없다는 오류가 발생하고 있습니다. quickstart.py에 있는 아래 코드를 수행시키면 from furiosa\_llm import LLM, SamplingParams llm = LLM.load\_artifact(“fur…

---

## [Llama 모델 build 및 logging 관련 질문](https://forums.furiosa.ai/t/llama-build-logging/356)

<div class="topic-metadata">

**Author:** [@gspark](https://forums.furiosa.ai/u/gspark)\
**Replies:** 10\
**Last updated:** [September 26, 2025, 11:58pm UTC](https://forums.furiosa.ai/t/llama-build-logging/356 "2025-09-26T23:58:35Z")

</div>

안녕하세요, 현재 퓨리오사 LLM 환경에서 Llama3.3-70B 모델을 빌드 및 실행하는 과정에서 몇 가지 이슈가 발생하여 이에 대한 질문을 드립니다. 병렬성 관련 퓨리오사에서 제공하는 Llama3.3-70B 모델은 TP가 32로 고정되어 있는 것으로 확인했습니다. 혹시 \*\*퓨리오사에서 제공하는 Hugging Face 기반 모델은 병렬성을 조절하는 커스텀 빌드\*\*가 가능한…

---

## [Throughput 성능 최적화 방법 문의](https://forums.furiosa.ai/t/throughput/362)

<div class="topic-metadata">

**Author:** [@gspark](https://forums.furiosa.ai/u/gspark)\
**Replies:** 1\
**Last updated:** [September 10, 2025, 5:26pm UTC](https://forums.furiosa.ai/t/throughput/362 "2025-09-10T17:26:39Z")

</div>

저희 모델(Llama3.1 모델을 post training한것)을 furiosa\_llm으로 빌드한것보다 Furiosa에서 이미 최적화하여 업로드 하신 동일 아키텍처 artifact가 30배 이상 token generation throughput 성능이 좋은 결과를 실험적으로 확인하였습니다. 이와 관련하여 동일 아키텍처 모델의 throughput 성능 최적화 방법에 대하여 문의드립니다. F…

---

## [More pre-compiled models including Qwen, QwQ and DeepSeek Distill in Hugging Face Hub](https://forums.furiosa.ai/t/more-pre-compiled-models-including-qwen-qwq-and-deepseek-distill-in-hugging-face-hub/354)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0\
**Last updated:** [August 27, 2025, 6:29am UTC](https://forums.furiosa.ai/t/more-pre-compiled-models-including-qwen-qwq-and-deepseek-distill-in-hugging-face-hub/354 "2025-08-27T06:29:25Z")

</div>

Hi folks, Additionally, we’ve uploaded the following models to Hugging Face Hub. furiosa-ai/QwQ-32B furiosa-ai/DeepSeek-R1-Distill-Qwen-7B furiosa-ai/DeepSeek-R1-Distill-Qwen-14B furiosa-ai/DeepSeek-R1-Distill-Qwen-32…

---

## [New Features: Structured Output and tool\_choice: "required"](https://forums.furiosa.ai/t/new-features-structured-output-and-tool-choice-required/352)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0\
**Last updated:** [August 26, 2025, 6:21am UTC](https://forums.furiosa.ai/t/new-features-structured-output-and-tool-choice-required/352 "2025-08-26T06:21:51Z")

</div>

Please check out the feature at the below document:

---

## [Llama 3.1 8b instruct 기반 sLLM 모델의 양자화 이슈](https://forums.furiosa.ai/t/llama-3-1-8b-instruct-sllm/333)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 7\
**Last updated:** [August 22, 2025, 8:47am UTC](https://forums.furiosa.ai/t/llama-3-1-8b-instruct-sllm/333 "2025-08-22T08:47:12Z")

</div>

안녕하세요? furiosa-llm (2025.3.0) 과 가이드 문서를 참고하여 llama 3.1 8b instruct 모델을 양자화하고 아티팩트를 성공적으로 만들고 추론까지 완료하였습니다. 그래서 저희 회사에서 llama 3.1 8b Instruct 모델을 pretrain 시켜서 Hugging Face에 공개한 unidocs/llama-3.1-8b-komedic-instruct 모델을…

---

## [Zstandard 라이브러리 오류](https://forums.furiosa.ai/t/zstandard/338)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 8\
**Last updated:** [August 14, 2025, 3:48am UTC](https://forums.furiosa.ai/t/zstandard/338 "2025-08-14T03:48:53Z")

</div>

안녕하세요? 계속 질문을 드리게 되네요.. 클라우드에서 인스턴스 생성시 스토리지 설정이 잘못되어 인스턴스를 재생성하고 RNGD에서 furiosa-llm을 테스트 중입니다. 아래와 같이 가이드에 있는 qickstart.py를 사용해서 인스턴스를 재생성하기 전에는 성공적으로 잘 수행하였습니다. from furiosa\_llm import LLM, SamplingParams -# Loa…

---

## [Fake quantize mode 이슈](https://forums.furiosa.ai/t/fake-quantize-mode/331)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 3\
**Last updated:** [August 7, 2025, 12:33am UTC](https://forums.furiosa.ai/t/fake-quantize-mode/331 "2025-08-07T00:33:45Z")

</div>

김종욱님 안녕하세요? 새로운 2025.3.0 버전으로 llm 모델을 성공적으로 로드하고 추론도 하였습니다. llama 3.1 8b instruct 모델로 가이드에 따라 양자화한 후 아티팩트로 만들어서 추론도 성공했는데요. 아티팩트로 만드는 과정에서 다음과 같은 에러가 발생하였는데 이 부분은 무시해도 되는 것인가요? “ERROR:2025-08-06 09:15:22+0000 Fake qu…

---

## [Furiosa-llm 모델 로드 이슈](https://forums.furiosa.ai/t/furiosa-llm/327)

<div class="topic-metadata">

**Author:** [@tobewiseys](https://forums.furiosa.ai/u/tobewiseys)\
**Replies:** 2\
**Last updated:** [August 6, 2025, 8:17am UTC](https://forums.furiosa.ai/t/furiosa-llm/327 "2025-08-06T08:17:04Z")

</div>

안녕하세요? GPU에서 개발된 sLLM을 NPU 버전으로 포팅하고 있습니다. 지난번에 RNGD 환경설정 이슈는 CSP의 지원으로 furiosa-llm이 정상적으로 설치되었습니다. furiosa-llm 2025.3.0 furiosa-llm-models 2025.3.0 furiosa-model-compressor 2025.3…

---

## [Offline Batch Inference Error after Upgrading SDK](https://forums.furiosa.ai/t/offline-batch-inference-error-after-upgrading-sdk/264)

<div class="topic-metadata">

**Author:** [@trick](https://forums.furiosa.ai/u/trick)\
**Replies:** 6\
**Last updated:** [May 19, 2025, 10:42pm UTC](https://forums.furiosa.ai/t/offline-batch-inference-error-after-upgrading-sdk/264 "2025-05-19T22:42:56Z")

</div>

After upgrading to Furiosa SDK 2025.1, I followed the documentation to build the model and run the offline batch inference script. However, I encountered the following error during execution: 2025-03-08T15:22:58.2286260…

---

## [Pre-optimized and pre-compield models in Hugging Face Hub](https://forums.furiosa.ai/t/pre-optimized-and-pre-compield-models-in-hugging-face-hub/278)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0\
**Last updated:** [May 19, 2025, 10:34pm UTC](https://forums.furiosa.ai/t/pre-optimized-and-pre-compield-models-in-hugging-face-hub/278 "2025-05-19T22:34:34Z")

</div>

Hi folks, Since the 2025.2 release, we’ve uploaded the pre-optimized and pre-compield models in Hugging Face Hub. For example, without an artifact compilation step, you can quickly run as follows: furiosa-llm serve fu…

---

## [An introduction to furiosa-llm (2025.1)](https://forums.furiosa.ai/t/an-introduction-to-furiosa-llm-2025-1/260)

<div class="topic-metadata">

**Author:** [@hyunsik](https://forums.furiosa.ai/u/hyunsik)\
**Replies:** 0\
**Last updated:** [March 3, 2025, 5:39pm UTC](https://forums.furiosa.ai/t/an-introduction-to-furiosa-llm-2025-1/260 "2025-03-03T17:39:47Z")

</div>
