Comment by UnboundedContex
3 hours ago
Approx ~131 million total. Made up of estimated:
102M - modified BERT-base-Chinese text encoder
26M - 3D U-Net-style vision/anatomy encoder
2.8M - projection layers, anatomy-specific projections, query tokens and attention layer
Written to run on something like an H100 though- as the CT Scan data is quite large.
No comments yet
Contribute on Hacker News ↗