Local LLM, Voice Agent & AI Infrastructure Lab
The problem
Practical conversational AI systems must balance privacy, response latency, speech quality, hardware constraints and operational reliability across several tightly coupled services.
What I built
I designed and operated a containerized lab integrating local large language models, GPU-accelerated Whisper speech recognition, LiveKit real-time communications, voice-agent orchestration and supporting automation. The work includes model and hardware evaluation, service integration, latency tuning, monitoring and deployment workflows.
Where it landed
The lab provides a working environment for experimenting with private AI assistants and real-time voice systems while demonstrating end-to-end capability across GPU infrastructure, speech pipelines, container operations, observability and production-minded AI integration.
