Research

Peer-reviewed work on AI/ML for cybersecurity, explainable and neuro-symbolic models, and the systems behind large-scale AI infrastructure.

Focus areas

  • AI infrastructure

    Multi-tenant GPU clouds on Kubernetes, scheduling and isolation, and the cost and memory trade-offs of LLM serving.

  • AI/ML for cybersecurity

    Intrusion and advanced persistent threat detection, and malware detection in virtualised cloud environments.

  • Explainable and neuro-symbolic models

    Interpretable deep learning for cyber-physical systems and neural-symbolic policy enforcement.

  • Distributed systems and observability

    Large-scale platform development, reliability engineering and observability at platform scale.

Peer-reviewed publications

Preprints

Working papers

LLM serving research in preparation. Links will be added as each paper is published.

  • GPU data center

    OpenDeepAPM: from request to silicon causality

    In collaboration with Anand Iyer and Mohammad Ali

    One unified trace per request: 27 spans across 7 layers.

  • Benchmarking

    OpenDeepBench: from request to silicon causality and workload-grounded benchmarking of LLM serving stacks

    In collaboration with Anand Iyer and Mohammad Ali

    Chat, RAG, agent, code and document-QA workloads scored end to end.

  • Routing & scheduling

    Learned routing for SLO-aware LLM serving

    In collaboration with Anand Iyer

    Prefill and decode pools, goodput, and P95 time to first token.

  • Speculative decoding

    Adaptive speculation under KV-cache compression

    In collaboration with Mohammad Ali

    Speculative decoding under 4-, 3- and 2-bit KV caches.

  • Memory & KV cache

    Tiered KV caching across HBM, DRAM and SSD

    In collaboration with Anand Iyer

    Session KV, predictive prefetch, and promote, demote and evict policies.

  • Cost–quality frontier

    Pareto Atlas: mapping the cost–quality frontier of LLM serving

    In collaboration with Mohammad Ali

    Calibrated simulation, the dominant configuration per regime, and the tensor-parallel versus KV-compression crossover.

Talks