>1B tokens/minute/GPU by combining query planner and inference engine

(modal.com)

7 points | by charles_irl 14 hours ago ago

2 comments