안녕하세요,
이전에 RNGD Hardware profiling 관련 문의글을 올렸는데, 추가로 확인해보고자 하는 정보들이 있어서 게시글 추가 작성드립니다.
RNGD Hardware profiling 문의 - 한국어 (Korean) / 일반 - FuriosaAI
- 아래 항목은 내부적으로는 측정이 가능할까요?
내부적으로 측정되지만, user 들에게 공개되지 않는 정보라면, 추후 API 또는 Tool을 open 할 예정인지 궁금합니다.
- Memory profiling
- Memory RW active time
- Device status
- Energy
- Communication
- PCIe active time
- PCIe transfer size (Byte)
- PCIe bandwidth
- 아래 항목들에 대해 task execution level 에서의 profiling이 가능할까요? 아래 문서를 참고하고 있습니다.
Furiosa-LLM 프로파일링 방법 - 한국어 (Korean) / 문서 및 튜토리얼 - FuriosaAI
- Memcpy / Memset task
- Calltree
- Memcpy / Memset start-end time
- Memcpy src / dst (PCIe)
- Memcpy src / dst (Internal)
- Memcpy execution count
- Memcpy / Memset size (Byte)
- Memory / Buffer allocation task
- Calltree
- Memory allocation func. active time
- Memory allocation size (Byte)
- Memory allocation address
- Memory 종류 (HBM or DM …)
- Communication task
- Calltree
- Synchronization start-end time
- NPU task
- Calltree
- ‘Furiosa-LLM 프로파일링 방법’ 튜토리얼의 방식에서 아래 기능도 지원이 되는지 궁금합니다.
- Marker - User가 Code 상에 marker 를 추가해 profiling json 에 추가
- Profiling 과정에서 유실 또는 오버플로 record count 확인
감사합니다.