找回密码
 立即注册
查看: 130|回复: 0

[求助]RTX A6000 fp16性能 和bf16性能?Tensor performance 性能是什么?

[复制链接]

737

主题

0

回帖

737

积分

00话痨hs000

积分
737
发表于 2023-7-10 16:56:16 | 显示全部楼层 |阅读模式
一、fp16性能  和bf16性能

GPU                 Compute Capability  来源于 https://developer.nvidia.com/cuda-gpus#compute  
RTX 6000                8.9                              
RTX A6000        8.6

根据 https://docs.nvidia.com/cuda/cud ... hmetic-instructions


Compute Capability   8.9  的fp16 性能= fp32性能
Compute Capability   8.6  的fp16 性能= fp32性能*2  《这不知道我理解的对不对》



二、Tensor performance 是什么?LLM 中训练用的性能与这个指标有关系吗?

                     
RTX A6000      Tensor performance      309.7 TFLOPS8   

NVIDIA A100   FP16 Tensor Core                 312 TFLOPS       采用稀疏技术的情况下


RTX 6000                Tensor performance    1457.0 TFLOPS      
这个简介里 “Fourth-Generation
Tensor Cores
Fourth-generation Tensor Cores provide faster AI compute performance, delivering more than 2X the performance of the previous generation. These new Tensor Cores support acceleration of the FP8 precision data type and provide independent floating-point and integer data paths to speed up execution of mixed floating point and integer calculations.”

是说这个只能是FP8的性能,就是用fp8推理时可以使用到的性能?  不能用于训练,哪怕性能减半也不能用于fp16训练?



NVIDIA H100   FP16 Tensor Core                 756 TFLOPS       采用稀疏技术的情况下
回复

使用道具 举报

您需要登录后才可以回帖 登录 | 立即注册

本版积分规则

Archiver|手机版|小黑屋|翻墙党

GMT+8, 2026-9-29 03:13 , Processed in 0.064096 second(s), 23 queries .

Powered by Discuz! X3.5

© 2001-2026 Discuz! Team.

快速回复 返回顶部 返回列表