Skip to content

7-8 OCT 2026​ | BENGALURU | INDIA

Pavel Sushin

Head of Analytics Tools Engineering & Operations, YTsaurus

About Talk

YTsaurus: Beyond the Hadoop Zoo — A Unified Platform for ETL, Analytics, and LLM Training at Exabyte Scale

Modern data platforms have become sprawling zoos — teams routinely run HDFS, YARN, ZooKeeper, Hive Metastore, HBase, Spark, Presto, and now Slurm or Kubeflow for GPUs, each with its own storage copy, security model, and failure surface. YTsaurus, open-sourced by Yandex under Apache 2.0 after 15 years in production, collapses that stack into a single distributed platform: one storage layer that serves batch ETL, sub-10ms KV lookups, ClickHouse and Spark analytics, and multi-node GPU training for LLMs — with unified access control, TLS, OAuth, and a Kubernetes operator for turnkey deployment. In this talk we tour the architecture (Cypress metadata, static and dynamic tables, GPU-aware scheduler), map each layer 1:1 to its Hadoop-stack equivalent, and show how the same cluster runs an exabyte-scale ETL pipeline, an interactive SQL workload via YQL, and LLM pre-training on shared hardware. You’ll leave with a concrete map of what YTsaurus would replace in your current stack — and a clear entry point if you want to evaluate it yourself.

TRACK: Keynote

8th Oct 2026 | Hall A | Time: 02:30-03:15 pm

About Speaker

Pavel Sushin is Head of Analytics Tools Engineering & Operations at YTsaurus, with over 10 years of experience as a C++ developer. He is a technical lead for the YTsaurus Platform, a large-scale storage and data-processing system, and specialises in distributed infrastructure and large-scale systems.