Research vllmarchitect
Optimizing Vision Encoders for Edge-Deployed Visual LLMs ICCCNet 2026 | published LNNS
- Benchmarked 8 architectures across 4 hardware platforms; TTFT ranging from 1.6ms to 98.0ms on identical hardware
- Identified vision encoders as the 70-85% latency bottleneck in edge VLLM deployments
- Introduced the Cache Cliff phenomenon and Sensory Bottleneck Hypothesis




