Making it easy to interface with Chalk’s query engine, with Chalk SQL and Chalk Dataframe

Bill Qin - Software Engineer
Chase Haddleton
by Bill Qin, Chase Haddleton
September 29, 2026

The data needed to make a prediction rarely lives in one place. Transactions might live in Postgres, historical data in BigQuery or Snowflake, events in a streaming system, and other context behind an API. To make a prediction, ML systems need to bring that data together, transform it into useful features, and make those features available at inference time. That means querying across different systems, applying the right logic, and doing it fast for production.

Historically, each part of this work has been done in different ways. Data scientists explore data with SQL or DataFrames. ML engineers build production pipelines in Python. Now LLMs and agents introduce even more ways to interface with data. AI teams are experimenting with new methods to provide agents access to fresh context, business logic, and data from the systems where it already lives.

The underlying problem is the same: how do you query, transform, and serve data across many systems while ensuring expressive and intuitive tooling for data teams (and agents) to query data? These are the hard problems that we built Chalk’s federated query engine to solve.

Having Chalk’s query engine speak the language of data scientists

Chalk is a federated query engine for context workloads across online, offline, and streaming environments. It queries data where it already lives, combines and transforms data across those sources, and executes the same logic at inference time.

Chalk SQL, Chalk DataFrame, and Python are different interfaces to the same underlying engine. With Chalk SQL and Chalk DataFrame, teams can use the interface that fits the task. The interface can change, but queries all run on the same execution engine that handles the work. The result is one system, with intuitive and expressive tools, for moving from data to decision.

A deep dive into Chalk SQL

We wanted to develop ways to work with our query engine in a way that data teams would find intuitive and expressive. We settled on Chalk SQL, a SQL dialect, and a dataframe library, chalkdf. SQL is declarative, widely understood, and the language for exploratory data analysis. It’s particularly well suited to describing transformations over structured data.

Chalk SQL brings that model to Chalk's federated environment. A single query can reference multiple connected sources and combine them into one computation. Instead of specifying every step the engine should take, a user describes the result they want and Chalk's planner compiles it into an optimized physical plan.

For example, a Chalk SQL query can join data from Postgres with data from BigQuery:

A federated query in Chalk SQL across a Postgres data source and a BigQuery data source

The interesting part of this query isn't the SQL itself. Rather, it’s that a user can describe the computation without deciding how data should move between our different sources. With the Chalk ecosystem captured in our catalog, Chalk SQL allows data scientists to answer any question about all their data.

Instead of the user first deciding how the data should be handled, the query engine decides. The engine receives a logical representation of the computation and can determine which portions of the work should execute where.

DataFrames give the same engine a different interface

SQL is a natural interface for querying data, but it isn't how every engineer thinks about a problem. Many workflows are inherently iterative including steps like loading data, filtering it, adding a derived column, joining another dataset, aggregating it, inspecting the result, changing the transformation and running it again.

DataFrames are designed around this style of work. Chalk DataFrame lets you work this way while still running on Chalk's engine and reaching every connected source. Rather than eagerly executing every operation as it is written, a sequence of DataFrame operations can be composed into a larger computation.

That distinction is important because the engine can optimize the computation as a whole. If an engineer writes a transformation that filters a large dataset and then selects a few columns, and every operation executes independently, then the system may materialize an intermediate result and move more data than necessary. Instead the operations remain part of one plan, and Chalk will reason about the complete computation. It can push eligible filters and projections toward the underlying source and avoid unnecessary intermediate materialization.

Since chalkdf is a dataframe library, it is very familiar to data engineers who use pandas and polars. This is particularly useful within Chalk-defined resolvers, which are Python functions designed to compute features from other features. Using chalkdf allows Chalk to compile these functions as the declared pipeline in the Dataframe instead of a blackbox Python function.

The user gets the programming model they expect from a DataFrame. The engine gets the information it needs to optimize the query. That is the pattern across Chalk's interfaces: familiar tools on top, with the same sophisticated query planning underneath.

Easily express external operations on the Chalk federated query engine

With SQL and Dataframe, you can also express various external sources plainly in these languages. The ability to integrate with external sources that also speak SQL is intuitive: querying a table from a Postgres source is expressed as querying a view in Chalk SQL. But even for sources that do not speak SQL it is possible to make those same expressive queries: both SQL and Dataframe can invoke functions that can hit TurboPuffer, Notion, Sagemaker, and OpenAI.

Recently, teams have been talking about models like Jev from Typesafe AI as a more cost-effective and faster alternative to traditional LLMs for generating structured decision outputs. Leveraging Chalk SQL, our co-founder Andy was able to easily add support for prompt_jev, allowing users to integrate directly with Jev within the Chalk ecosystem.

To demonstrate, consider the following Chalk SQL query.

SELECT
  id, 
  prompt_jev(prompt, ‘{“u”: {“type”: “noul”, “instructions”: “friendly”}}’) AS friendly_score,
FROM my_snowflake.user_prompts 
WHERE user_id = 12345 
LIMIT 100

This query extracts up to 100 prompts from a Snowflake source for a specific user, then runs Jev on the prompts to determine how friendly each one is. When federating with external sources, it’s important to be able to push down work into those sources. Since this executes within Chalk’s query engine, if this query were to execute without pushdown the plan would look something like the following unoptimized path.

This would result in a sequential scan of the snowflake table user_prompts. So, we have to parse the above statement and pushdown the work. In this case, we can have Snowflake apply both the user_id filter and the LIMIT clause, and then apply Jev to those results.

How a query would execute without or with pushdown

This was just a simple example of pushdown, but this pattern gets increasingly difficult when the filters and projections used get more complex. For example, notice we could not push down the prompt_jev itself to Snowflake: Snowflake doesn’t have an equivalent function. But, Chalk can recognize that it does need any other columns from the user_prompts table other than id and prompt, since those are the only two columns referenced in the output. These optimizations are critical in ensuring that SQL and Dataframe feel good to use and are the most efficient they can be.

These tools make it easy to interface with Chalk’s unified query engine 

We added Chalk SQL and Chalk DataFrame as two new frontends to Chalk’s query engine. Both translate queries into our existing logical representation, bringing our years of feature-serving optimizations to these interfaces and allowing SQL and DataFrame operations to be composed into existing Online and Offline queries. From this shared representation, Chalk’s logical optimizer runs a multi-pass rewrite loop over the plan, simplifying and fusing the computation before the execution framework lowers it into an executable plan.

Chalk’s logical optimizer runs a multi-pass rewrite loop over the plan, simplifying and fusing the computation before the execution framework lowers it into an executable plan.

Behind the scenes, Chalk’s execution framework primarily uses Velox, Meta’s open-source, composable C++ execution engine. Velox provides building blocks to support running large-scale analytical queries out-of-the-box, but Chalk has a unique set of requirements for our query engine that differ from a classical analytical engine. We target both ends of the low-latency <> high-throughput spectrum, so we have customized and optimized Velox to serve both of our needs.

To date, we have introduced 26 custom Velox operators including:

  • Zero Copy HashJoin
  • IO Expressions
  • User Defined Functions
  • As Of Joins

These custom operators both allow us to optimize our query performance, but also integrate our engine into 10+ databases. Our performance-focused commits mean that Chalk can execute and serve queries in single-digit milliseconds at p99.9 under substantial load. So, from the user’s Chalk SQL, DataFrame operations, or feature definitions, workloads share the same underlying execution model.

One engine from data to decision

The goal of Chalk SQL and Chalk DataFrame is to give data and AI teams familiar ways to express computations that run on Chalk’s federated query engine. Chalk SQL provides a declarative interface for querying and transforming data. Chalk DataFrames provide a composable interface for iterative analysis.

Federation lets those computations reach data where it already lives. The planner determines how the computation should be executed. Velox provides the underlying execution engine for the resulting plan. Chalk also compiles supported Python logic into Velox expressions, allowing Python-authored computations to benefit from native, vectorized execution. You can learn more about how Chalk statically compiles Python functions in Nathan’s blog post.

The computation is handled by the Chalk query engine. So, different ways of expressing a computation doesn’t require different systems to execute it. When the interface, planner, federation layer, and execution engine work together, the path from exploration to production gets shorter.

Explore Chalk SQL →

Learn how to construct a Chalk DataFrame →

Run experiments in Chalk Notebooks →

Want to stay up-to-date with Chalk?

Subscribe for updates on what we’re building (and shipping!) at Chalk