BRIEFING
ONNX로 Transformer를 내보낼 때 생기는 함정
ONNX로 Transformer를 내보낼 때 생기는 함정
Scope
핵심 질문
Transformer 모델을 ONNX로 export할 때 왜 opset, dynamic shape, attention 연산과 cache 때문에 문제가 자주 생기는가?
Research
핵심 근거
- ONNX portability는 source framework 연산이 표준 operator 또는 지원되는 custom operator로 변환될 수 있어야 한다.
- opset version과 target runtime의 operator support가 맞아야 한다.
- Transformer는 dynamic sequence length, attention mask, KV cache, control flow 등 때문에 단순 CNN보다 export 조건이 복잡하다.
참고 자료
- https://onnx.ai/onnx/repo-docs/IR.html
- https://onnx.ai/onnx/intro/
Draft
Transformer를 ONNX로 내보내는 일은 model.save 형식 변환보다 계산 그래프 호환성 문제에 가깝다.
Opset
Exporter가 생성하는 operator semantics와 target runtime이 지원하는 opset이 맞아야 한다. 너무 최신 opset은 runtime이 지원하지 않을 수 있고, 너무 오래된 opset은 필요한 연산이 없을 수 있다.
Dynamic Shape
Batch와 sequence length가 실행마다 바뀌면 input dimension을 dynamic으로 표현해야 한다. Export 예제에서 고정 길이로 trace하면 실제 serving에서 다른 길이를 넣을 때 실패할 수 있다.
Attention
Framework가 fused attention이나 custom kernel을 사용하면 표준 ONNX operator 집합으로 바로 표현되지 않을 수 있다. Exporter가 여러 primitive op로 풀거나 custom op를 남겨야 한다.
KV Cache
Autoregressive generation에서는 past key/value를 입력으로 받고 새 cache를 출력하는 graph가 필요하다. 단순한 full-sequence forward graph만 export하면 효율적인 decoding에 쓰기 어렵다.
검증
Export가 성공했다는 것과 결과가 맞다는 것은 다르다. 여러 shape와 입력으로 source framework 결과와 ONNX runtime 결과를 비교해야 한다.
핵심
Transformer ONNX export의 핵심은 opset, dynamic shape, attention implementation과 cache interface를 target runtime 기준으로 맞추는 것이다.