
0:00
53:20
On episode 4 of Lab Notes, Amir Zohrenejad speaks with Junchen Jiang about why KVCache may be better understood as reusable, AI-native data rather than a temporary inference optimization. They explore how LMCache and CacheBlend can reduce redundant computation, move context across distributed inference systems, and help support increasingly complex AI agents. The conversation also covers multimodal workloads, open-source infrastructure, and the future of AI systems research.
More episodes from "Heavybit Podcasts"



Don't miss an episode of “Heavybit Podcasts” and subscribe to it in the GetPodcast app.








