
0:00
53:20
On episode 4 of Lab Notes, Amir Zohrenejad speaks with Junchen Jiang about why KVCache may be better understood as reusable, AI-native data rather than a temporary inference optimization. They explore how LMCache and CacheBlend can reduce redundant computation, move context across distributed inference systems, and help support increasingly complex AI agents. The conversation also covers multimodal workloads, open-source infrastructure, and the future of AI systems research.
D'autres épisodes de "Heavybit Podcasts"



Ne ratez aucun épisode de “Heavybit Podcasts” et abonnez-vous gratuitement à ce podcast dans l'application GetPodcast.








