papersTODAY 04:00 UTC
SpliTEE combines trusted hardware with differentially private GPU offloading for LLM inference
A new paper proposes SpliTEE, a system that runs large language model inference partly on trusted hardware while outsourcing the rest to GPUs with differential privacy guarantees. The approach aims to keep user prompts confidential, addressing risks such as sensitive data being memorized during retraining by remote model providers. It targets a balance between privacy protection and inference performance.