Hyperspace - An Indexing Subsystem for Apache Spark

Overview

Explore the design, implementation, and operationalization of Hyperspace, an indexing subsystem for Apache Spark, in this 32-minute conference talk by Databricks. Learn about the foundations of the indexing infrastructure, including API design and integration with Spark's Catalyst optimizer. Discover how Hyperspace enables users to build, maintain, and leverage indexes on various data formats for query acceleration and resource cost reduction. Gain insights into the multi-user concurrency model and the development roadmap for open-sourcing this technology. Through presentations, benchmarks, code examples, and notebooks, delve into the world of efficient data indexing for large-scale datasets ranging from GBs to PBs, addressing both batch-style queries and explorative analytics.

Syllabus

Introduction
Who are we
What is an index
Overview
Investment
APIs
Index Creation
Index Benefits
Demo
Investment Areas
Hyperspace Types

Taught by

Databricks

Reviews

Start your review of Hyperspace - An Indexing Subsystem for Apache Spark

Taught by

Care and Feeding of Catalyst Optimizer - Practical Troubleshooting for Spark SQL

Deep Dive into GPU Support in Apache Spark 3.x - Accelerator-Aware Scheduling and RAPIDS Plugin

Koalas: Scaling Pandas APIs on Apache Spark - Performance and Comparison with Dask

Optimizing Catalyst Optimizer for Complex Spark Plans

How Apache Spark 3.0 and Delta Lake Enhance Data Lake Reliability

User Defined Aggregation in Apache Spark - From Challenges to Improvements

Never Stop Learning.