Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon

(github.com)

166 points | by tatef 5 hours ago ago

54 comments