# Homepage source: https://chalk.ai/ Chalk | The data platform for AI + ML Chalk is the real-time data platform for ML & GenAI. More than a feature store, Chalk delivers fresh features in under 5ms with on-demand compute and caching in your cloud. The data platform for AI + ML Ultra-fast data pipelines, caching, on-demand compute. All in your cloud. Announcing our $50M Series A Customers - Mission Lane - Melio - Whatnot - Socure - Iwoca - Moneylion - Apartment List - Grindr - Pipe - Turo Feature pipelines in idiomatic Python Powerful data engineering workflows, without the infrastructure headaches. Powered by Rust. Built-in scheduling, streaming + caching Complex streaming, scheduling, and caching, all defined in simple, composable Python. Composed & queried in real-time Make ETL a thing of the past. Fetch all of your data in real-time, no matter how complex. Toolchain for LLMs Incorporate deep learning and LLMs into decisions alongside structured business data. Machine learning infrastructure is painful. Chalk makes it simple for data teams to focus on building the unique products and models that make their businesses thrive. Customer Testimonials We’re moving hundreds of millions of features per second, each payload around 1MB, and still hitting a P99 latency of just 100ms. That kind of performance across the board is a real testament to the system Chalk built. - Emmanuel Fuentes, VP, Data & AI, Whatnot Chalk helps us deliver financial products that are more responsive, more personalized, and more secure for millions of users. It’s a direct line from infrastructure to impact. - Meng Xin Loh, Technical PM, Moneylion Chalk has become a powerful addition to our machine learning infrastructure at Mission Lane. Chalk has enabled us to unify and streamline our feature calculations across both offline/batch-eval and online/live-decisioning use cases. We continue to be impressed by the flexibility and scalability of the system, and by the willingness of the Chalk team to work with us to get even more value out of it. - Mike Kuhlen, Data Science & Machine Learning Solutions and Strategy, Mission Lane We’re applying AI and ML at scale across key areas of our energy business with Chalk’s feature platform. It enables high-performance computation over diverse data sources using clean, reusable code. The ability to mix Python and SQL gives our team the flexibility we need, while shared feature logic across projects improves consistency and accelerates development. - Edward Li, Staff AI/ML Engineer, Sunrun By moving our feature pipelines to Chalk, Data Science and Engineering now work side-by-side throughout model development. What used to be lengthy, error-prone handoffs are gone. Our entire search ranking stack runs on Chalk, serving features for inference in under 50ms, and we’re extending it to all of our models, including real‑time personalization. - Moaj Musthag, Head of Engineering, Turo Chalk powers our LLM pipeline, turning complex inputs—HTML, URLs, screenshots—into structured, auditable features. It lets us serve lightweight heuristics up front and rich LLM reasoning deeper in the stack, so we catch threats others miss without compromising speed or precision. - Rahul Madduluri, CTO, Doppel Our competitive edge hinges on the speed at which we analyze risk patterns, test new detection features, and deploy updates. With Chalk, our iteration speed went from days to hours. - Raine Scott, Co-Founder, Verisoul Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project. What would have taken months took weeks. - AJ Balance, Chief Product Officer, Grindr Deploy to your own infrastructure: Use your existing database as your online + offline store. No bespoke storage. Everything in your cloud. High-volume workloads at ultra-low latency: Chalk's Compute engine scales horizontally out-of-the-box and executes most complex queries on a Rust-based runtime for maximum performance. 100,000 QPS in < 5ms? We have you covered. Power real-time decisions with real-time data. Make better predictions with fresher data. Don't pay vendors to pre-fetch data you don't use. Query data just-in-time for online predictions. Perfect auditability. Know everything you computed and data replay anything. Unify training and serving. Iterate faster. Experiment in Jupyter, then deploy to production. Prevent train-serve skew and create new data workflows in milliseconds. Detect, troubleshoot, and eliminate data issues. Track data use, drift, and quality effortlessly with observability—built right in. Integrations Integrate with the tools you already use and deploy to your own infrastructure. - PostgreSQL - Snowflake - AWS - Google Cloud - Databricks - Jupyter - Datadog - PagerDuty - Slack - Python - Apache Arrow - Apache Airflow Start building with Chalk: https://chalk.ai/book-demo Helpful homepage links - Book a demo: https://chalk.ai/book-demo - Documentation: https://docs.chalk.ai/docs/what-is-chalk - Code Examples: https://chalk.ai/code-examples - Chalk Blog: https://chalk.ai/blog - Customer stories: https://chalk.ai/customers - Careers at Chalk: https://chalk.ai/careers - About Chalk: https://chalk.ai/about # About source: https://chalk.ai/about Data infrastructure to answer the world's hardest questions Chalk's data platform provides the essential building blocks for machine learning with an experience developers love. Why Chalk? Solving large-scale ML problems is a difficult problem that requires mathematical precision and a lot of hard work. Our name is inspired by the favorite tool of mathematicians around the world—the chalkboard. Our logo is taken from the QED symbol, which indicates the end of a mathematical proof. Founders With a combined 35 years of experience working in industry, our founders have solved some of the most difficult problems in finance, data infrastructure, and risk management at companies like Google, Stripe, Affirm, and Palantir. Marc Freed-Finnegan - Co-Founder & CEO Marc spent many years at Google where he helped to launch the first version of Google Wallet. He went on to start Index, which Stripe acquired as its in-store payment solution—now called Stripe Terminal. Elliot Marx - Co-Founder Elliot started his career at Affirm where he built the early risk and credit data infrastructure system (the inspiration for Chalk). He then co-founded Haven Money, which Credit Karma acquired to power its banking products. Andy Moreland - Co-Founder Andy worked at Palantir on large government data infrastructure projects. He then co-founded Haven Money (with Elliot), which now powers Credit Karma Money. Careers Work with some of the world's top engineers on cutting-edge problems in systems engineering, query planning and optimization, and distributed analytical data processing systems. We are a diverse group of people working to change the world of machine learning and bring powerful new tools to the industry. # Careers at Chalk source: https://chalk.ai/careers Careers at Chalk Work with some of the world's top talent on cutting-edge problems. We're creating the data platform that fundamentally changes what's possible for developers. Why our team loves working at Chalk Working here pushes you to master systems programming at a speed you can't get anywhere else, because the problems we tackle here have a theoretical bent that's genuinely hard to find in corporate software jobs. This is where you come to level up. - Sai Atmakuri, Software Engineer When I first came to the Chalk office for my onsite, I saw an interesting problem scribbled on the chalkboard. When I asked how they solved it and they mentioned bloom filters, I knew right away that these were the people I wanted to build with. - Bill Qin, Software Engineer What began as wild whiteboarding sessions in our NY office has evolved into a real growth engine. We're building a GTM team that's sharp, creative, while doubling down on what works. If you want to help shape how truly groundbreaking AI reaches the market, this is the place to be. - Alexandra Kane, VP of Revenue I think being around sharp, driven people naturally elevates your own work. It's no surprise we have chess IMs and competitive Smash players on the team - people here bring that same intensity to everything they do. It makes our weekly board game nights pretty memorable too. - Kelvin Lu, Software Engineer If I had to describe Chalk in a nutshell, it's nice nerds solving interesting problems together. One of my favorite activities here is the biweekly book club on AI/ML - it's refreshing to collaborate with people who are brilliant but also down-to-earth and fun to be around. - Melanie Chen, Forward Deployed Engineer It's cool to get a front row seat to a company powering the backbone of one of the most exciting industries in the world right now. But what truly keeps me here is the culture of ownership. No BS, no office politics, just the freedom to move fast and see your impact immediately. - Brian Sussman, Operations Manager Life at Chalk You'll find us celebrating wins together — whether that's solving complex problems or closing important deals — while battling in board games and going on culinary adventures to fuel our energy. Benefits you'll receive - Full medical, dental, and vision insurance - Flexible Spending Account (FSA) and Health Savings Account (HSA) - Expert healthcare guidance and live one-on-one support - 401(k) retirement savings - Generous holiday schedule - Generous PTO annually - Daily lunch and dinner on us - Flex commuter benefits - Grab a ride home on us # Chalk Notebooks: where your ML agent does its work | Chalk source: https://chalk.ai/blog/announcing-chalk-notebooks A workspace for agentic ML. Your agent investigates against the data your models actually see, and the notebook it leaves behind is the evidence you review before deciding what ships. Today we're announcing Chalk Notebooks, built for agentic ML. Shipping a model is an exercise in confidence: that the model and features you trained on will behave the same way in the real world, and that the real world will keep behaving like the world you trained on. Agents, even on frontier models, test that confidence. They rapidly analyze data and write code, but they can miss basic steps, fumble judgment calls, and hand back overconfident answers. Ask one if it's sure and it might fold or double down, with little in between. When the cost of being wrong is high (and in production ML it is), trusting agents on their word is professional negligence. The chat log holds the agent's side of the story, but it can't answer the question a model change turns on: how will production behave with the new model? No test proves an analysis the way tests prove code, so validating a model change means following the reasoning and evidence behind it. Short of a randomized controlled experiment on live decisions, the best you can do is replay history using only what each decision knew at the time. Whether you are in the loop, following the agent's steps, or on the loop, reviewing them afterward, the reasoning and evidence must outlive the chat session. That is reproducibility, and in ML it starts with the data. Why we built Chalk Notebooks Production ML is full of subtle sharp edges, and accidentally analyzing data that doesn't represent production is one of the sharpest. We built Chalk Notebooks to take that edge away. Chalk assembles data from your production sources and serves it fresh to your models. A Chalk Notebook is a hosted notebook whose code runs inside your own cloud, next to your production deployment. It queries that deployment directly, so you and your agents record and reproduce analysis on the data your models actually see. Chalk Notebooks have Python, SQL and markdown cells, a chalk notebook command-line interface (CLI) and a Model Context Protocol (MCP) server, so an agent can drive them. They also support three things a notebook off the platform cannot: - One live federated query across your application databases, warehouses, streams and APIs, plus the online and offline features built on them - Point-in-time correctness as a query parameter - Production and a Chalk branch (a copy of your deployment serving no production traffic), queried from the same cell Watch it work In this demo, my agent will use a Chalk Notebook to run federated SQL queries across production databases to find the driver of a loan default spike, build point-in-time features, retrain the model on Chalk Model Training, test the candidate on a Chalk branch against production, and hand me a candidate model for promotion and the analysis that backs it up. Say I run the loan approval model for a point-of-sale lending startup. Over Sunday coffee I see a streak of bad loans: first-payment defaults are up sharply over two months and my model never flagged it. I open a connection to my coding agent running on a Chalk sandbox in my production environment. My sandbox has my model's repo mounted as a Chalk volume, the chalk CLI installed, and Chalk's MCP server connected. I ask it to investigate the issue. The agent reads the repo and runs a few orientation queries via the Chalk MCP Server. After getting oriented to the problem, it creates a Chalk Notebook via the Chalk CLI. I open the notebook it created in the Chalk dashboard and can see it start adding cells in real-time. But knowing the agent's investigation may take a while, I decide I don't want to be a human in the loop. I finish my cortado, shut my laptop and walk home. Later, the agent notifies me that the notebook is complete, including the analysis and style peer reviews. I start my review. (Feel free to explore the agentic ML notebook on your own here.) Two things jump out from the summary. First, a Florida pawn shop fraud ring may be involved. Second, it got there using most of the platform, and the notebook shows its work. What happened. The notebook's exploration starts with queries to understand the problem. Below, you can see how the agent used a federated SQL query to analyze weekly loan decision data and outcomes across the online lending_db and the offline analytic data warehouse (the join in blue). The accompanying charts show that the funded default rate had been at 2.0-2.4% from March to May, but jumped to 3.4-3.8% starting in June, while the model's mean prediction remained around 2.0%. Who and where. The agent slices the default increase by merchant category, geography and merchant age, and one slice holds the key. Six Florida pawn shops onboarded this year hold 355 of the 435 pawn-shop defaults and together have a default rate of 90.8%, while every other merchant sits at 2.5%. Worse, my production model predicted a default rate of 1.91% for them, safer than the rest of the loan book. The signal appears to be related to the merchant, but my model was only looking at signals on the borrower. The features. The agent identifies six point-in-time features about the merchant and its borrowers and builds a training frame of 72,837 loans in one ChalkSQL query, each computed as of that loan's own origination. The leak rules are written into the query predicates: a feature can use only what was knowable when the loan was approved. The resolvers that later compute and serve these features in real-time carry the same predicates, so training and serving agree by definition. The agent validates that 300 sampled loans on the Chalk branch all match the training data frame. It shows me how these features differ between the six Florida pawn shops and all other merchants. The models. Every model fit runs on Chalk Model Training. On a holdout of 12,242 labeled loans, set aside because they originated in the four weeks after the training window, and at a matched 95% approval rate, the model candidate cuts the default rate among approved loans from 3.4% to 2.7%. However, the agent sees that retraining on production's own six inputs with fresher training data gets down to 2.9% without new features, so most of the gain is from the fresher training data. The agent realizes that, because loan outcomes aren't discovered for over a month after the loan is originated, the training window actually contains information about the six Florida pawn shops that the model couldn't have known when the defaults started. So the agent runs a more honest walk-forward analysis that retrains the model every week on only the outcomes reported by that Monday. That candidate declines practically nothing from the six merchants for six weeks, then declines about two thirds of the six merchants' defaults. The agent finds an additional rule about the merchant on top of the model that catches all of the segment's defaults and fires on no other merchant's loans, bringing the default rate down to around 2%. Checking against production. The agent then runs a live test between production and the Chalk branch. chalk apply --branch deploys the features, resolvers, and artifact beside production. On the last four weeks' worth of applications, production approved 189 of the six merchants' 190. The Chalk branch approves none, worth $364,390 of principal, and approves 96.7% of an 1,800-loan sample of everyone else against production's 95.8%. Asked live, the way the approval service asks production, one borrower's $1,900 loan application at one of the six Florida pawn shops scores a loan default prediction of 3.1% on production (approved) and 73.1% on the Chalk branch (declined). The same borrower at an established merchant scores 3.0% on the Chalk branch (approved). The blue highlight below is the online query the agent used for the analysis. The recommendation. The agent recommends promoting three things in this order of confidence: the merchant rule, the retrained model, and the six features. (I'm also left with the questions about how to handle the six Florida pawn shops, and how to treat new merchants with no reported outcomes.) Looks like I'll have a productive Monday. Governed by default... An agent with your production data is a different security posture than an engineer with it. If you skip permissions on a React frontend, you might lose an afternoon, but if you skip them on your production data, you might lose the data. This one held a service token scoped to one environment, called its model through the Chalk Model Gateway on a budgeted, revocable key, and worked in a sandbox that mounted only the repo it needed. Write access is granted per environment: on in dev, off in prod. Row and column access policies apply to agent-generated SQL, and Chalk Notebook kernels are sandboxed compute with egress you control. More on governance and observability. ...with a human on the loop Everything above happened in one notebook against the deployment that production runs on: federated queries found the segment, point-in-time features described it, Model Training fit the candidate, and a Chalk branch scored it against production. The agent was never able to change production. This is as it should be. The agent investigates; the human ships. Get started Chalk Notebooks are available now. Open one, connect your coding agent to the Chalk MCP server, and get started with agentic ML on Chalk. # Agentic Machine Learning on Chalk: Introducing MCP Server and Chalk Assistant source: https://chalk.ai/blog/announcing-chalk-mcp-server Chalk MCP Server connects AI agents directly to the Chalk primitives your team already uses and Chalk Assistant brings that agent inside the Chalk interface. Every engineering team is bringing AI agents into the way they work. But for most data and ML teams, those agents still sit outside the systems and workflows where model development actually happens. Agents wait for context to be pasted into a prompt, then hand back suggestions that need to be manually tested in the real system. That’s useful, but it doesn’t address the bottleneck. The most time-intensive work for data scientists and ML engineers is gathering context, running investigations and turning findings into new features, experiments, and model versions. Today, we're announcing Chalk MCP Server, which connects AI agents directly to the capabilities you use in Chalk to build, evaluate, and deploy ML systems, and Chalk Assistant, which brings an agent inside the Chalk product experience. Together, Chalk MCP Server and Chalk Assistant provide the foundation for agentic machine learning: agents that don’t just help from the sidelines, but participate in the actual loop of developing ML systems. Connect agents to the systems where ML work happens The Model Context Protocol (MCP) is the standard for connecting AI applications to external systems. Chalk MCP Server exposes the capabilities engineers already use in Chalk directly to AI agents. Through Chalk MCP Server, an agent can: - Read feature and resolver definitions and suggest targeted changes. - Run online queries against Chalk features and resolvers. - Use Chalk SQL for feature research, extraction, and analysis. - Engineer and backtest new features to improve model performance. - Inspect query errors and behavior, including performance characteristics. - Explore data and features, and support feature engineering and model investigation workflows. This gives an agent access to the same primitives you use to investigate a model, identify an opportunity, build a feature, backtest, and iterate. Let agents drive the improvement loop With Chalk, ML teams cut their feature and model development cycles from weeks to days. Chalk MCP Server accelerates that loop even further by helping agents investigate, validate, and ship changes. With agents connected to Chalk, teams unlock faster model iteration and proactive improvement. Instead of waiting for incidents or in review queues, agents continuously test changes. The acceleration compounds: more experiments, tighter feedback loops, and models that get better before problems surface. Chalk Assistant: the in-product agent experience Chalk Assistant is how that connection shows up in the product. In a simple interface, you can: - Bring an agent into Chalk with your own model API key. - Ask the agent in a sidebar to investigate a model, feature, or query using Chalk’s context and tools. - Review the agent’s reasoning, queries, and proposed changes without leaving the interface. Chalk already gives you the tools for building, evaluating, and deploying AI systems, solving common challenges like unifying training and serving, building point-in-time correct datasets, and providing observability across queries and features. Chalk Assistant puts an embedded collaborator inside your ML workflows, right where you already work. Move from AI help to agentic machine learning The goal is to give agents the ability to participate in the work, not just comment on it. An agent connected to the systems where your data, logic, and models live can perform the improvement loop directly. The future of ML development is agents with the context, access, and controls to run the full loop of building and operating production ML systems. Get started Chalk MCP Server and Chalk Assistant are available now. Install the MCP Server to connect your agent to Chalk, or reach out to your Chalk team to learn more. # What is agentic machine learning? | Chalk source: https://chalk.ai/blog/agentic-machine-learning Agents now run the modeling loop end to end. What decides whether you can trust it is the data layer underneath. TL;DR. Agentic machine learning is agents running the modeling loop itself: propose a hypothesis, implement it, evaluate it against a metric, keep it or discard it, repeat. A person still sets the objective. The agent optimizes toward it, completing dozens of cycles overnight instead of a handful a week. The loop is not the hard part to build. What decides whether a team can trust it is the data layer underneath: whether the agent sees what production saw, and whether the feature it proposes can be served without a rewrite. A pricing model has been leaving margin on the table since Friday. Someone has to work out why on Monday morning. Investigating it is manual and slow. An engineer has to reconstruct what the model saw at the moment it scored, then compare that against the training data. Finishing in a day is a good outcome. Now multiply that by every missed prediction worth chasing. Teams don't need to devote an engineer to that work. An agent can take the first pass. It pulls the relevant decisions, replays the context each one ran on, and drafts a candidate fix. An engineer evaluates the work and ships the change, so a day-long effort becomes a few minutes of review. That loop has a name now. Agentic machine learning is what happens when agents run the modeling loop. What follows is what the loop requires to be trusted, and that turns out to have less to do with the agent than with the data underneath it. How does the agentic machine learning loop work? Agentic machine learning hands the model development loop to an agent. The agent proposes a change, implements it, runs it, scores it against a metric, and keeps or discards it based on the result. A human sets the objective and the agent optimizes toward it. The loop has six steps: 1. Propose a hypothesis. 1. Implement it in code. 1. Execute it and produce a metric. 1. Evaluate whether the result beat the baseline. 1. Commit the change or roll it back. 1. Repeat, building on what was kept. It's the loop ML engineers have always run. Context engineering is a new name for feature engineering, which is a new name for data engineering. The name changes and the job does not: find the signals that drive decisions, get them to the model, and make sure they mean the same thing in training that they mean in production. What changed is the rate. A person runs a handful of model iteration experiments a week. An agent runs dozens overnight. What do teams gain from agentic machine learning? Teams gain throughput on work they already know they should be doing and cannot reach. The failures you currently let through. Every team triages: the loudest three misses get investigated and the rest go on a list nobody returns to. An uninvestigated miss still charges the business every time it fires. A team that can work the whole queue instead of the top of it stops paying for the same mistake twice. Iteration speed. Model quality is a function of how many cycles a team completes, and most teams are capped on cycles by how long the hand-work takes rather than by ideas. Compressing the investigation means the same headcount runs more experiments against live results. Headroom across the organization. ML teams are asked to support more use cases each year than they gain headcount for. Every new model brings its own investigation load and its own failures to chase. Teams stall when that load scales linearly with the number of models in production. Handing the repetitive layer to agents is one of the few ways to take on new use cases without growing the team. How is agentic ML different from AutoML, coding agents, and recursive self-improvement? AutoML searches a space you defined in advance, while an agent can redefine the space. A coding agent's output either compiles or it doesn't, while an ML agent's output is a number that looks the same whether it is right or wrong. Recursive self-improvement is the same loop taken further, where every accepted change becomes the baseline the next experiment has to beat. Agentic ML vs AutoML AutoML searches a space you defined in advance. You pick the model families and the ranges. The search explores inside those walls, and it has been good at that for years. An agent changes the walls. It can decide the problem was framed wrong, redefine a label, or introduce a feature nobody had considered. AutoML optimizes within your hypothesis. An agent generates hypotheses. Agentic ML vs coding agents A coding agent ships code that either compiles or does not. Tests pass or they fail. The feedback is close to binary and it arrives fast. An agent doing machine learning ships a number. The number looks the same whether it is right or wrong. That difference is why general-purpose coding infrastructure does not transfer cleanly to this work. The code compiling tells you almost nothing about whether the result is real. Agentic ML vs recursive self-improvement Recursive self-improvement sits at the further end of the same spectrum. Agentic ML describes agents running the loop. Recursive self-improvement describes a loop that compounds, where each accepted change becomes the baseline the next experiment has to beat. Even Karpathy's autoresearch has an agent optimizing a different model rather than itself. The tool hands an agent a small language model training setup, runs a fixed five-minute experiment, keeps or discards the change on a single metric, and repeats. The README puts throughput at roughly 12 experiments an hour, or about a hundred while you sleep. In the run he reported in March, an agent ran about 700 experiments over two days and found 20 optimizations. Applying those 20 changes to a larger model cut training time by 11 percent. The agent tunes the training code and initial settings for a different, smaller model. It never touches its own weights. What is shipping today is a bounded loop with a human gate on promotion, moving toward fewer gates as evaluation gets more trustworthy. Does agentic ML work on mature models, and how does it fail? Yes, including on models teams had already spent years tuning. The failures are data problems. Instacart published the most detailed account so far. Their delivery time model is a traffic-aware gradient-boosted tree with years of iteration behind it, and they were, in their words, genuinely skeptical that an autonomous research loop could improve it. Several engineers ran independent loops against the same data snapshot, the same splits, and the same validation metric. A 30-trial LightGBM search cut held-out mean absolute error by 3.6 percent against production. A tuned-MLP search cut it 4.8 percent against its own tree-based baseline. On catalog attribute extraction, a 25-round loop improved recall by 8.1 points with precision held above its floor, replacing a process that had taken about a week per attribute. The gains stacked up out of small wins. Most of the agent's hypotheses failed. Their write-up also lists what went wrong, which is helpful background because most teams only publish the wins. - A point-in-time join error that merged the future onto the present. - Feature leakage, recurrent, and described as particularly sinister because it is so hard to spot in offline analysis. - An agent reading the held-out evaluation set while it selected hyperparameters. - Features that added latency, or that needed data production does not have. In every one of these the metric went up. A join that pulls the future into a training row produces a feature that predicts the label. So does a leaked one. An agent that has seen the held-out set scores well on it. The loop keeps whatever improves the metric. That is the whole mechanism, and there is nothing in it to catch any of this. At four experiments a week an engineer reads each one, and a feature correlating a little too well with the label gets noticed. At four an hour nobody is reading. The loop also prefers these features, because a leaked one wins. By morning it has twenty more experiments built on top of it. How often should engineers review what an agent produces? Often enough that no change reaches a production model without a human approving it, and rarely enough that review does not become the new bottleneck. Handing work to an agent mirrors a process your team already runs. The agent opens a pull request. It proposes a change and shows the work behind it. An engineer reviews, comments, and merges or rejects. Engineers can review far more proposals than they can write. That moves a team's ceiling from how fast its people type to how fast they can judge. Changes that reach a production model still get human approval. Narrow, reversible changes with nothing downstream depending on them can run behind a shadow evaluation on their own, so review effort concentrates where the risk actually sits. An agent may run unattended overnight and produce forty accepted changes by morning. Reviewing each one recreates the bottleneck. Reviewing none of them means trusting a metric that looks identical whether it is right or wrong. How do you set up evaluations for agentic machine learning? Start by deciding what you are evaluating, because there are two answers and they need different machinery. One is the model change the agent proposes, scored against a point-in-time correct backtest. The other is the agent's own behavior, meaning whether its plan and tool calls were reasonable and what they touched along the way. A loop that only measures the first will accept a good number produced by a bad process. For the model change, the mechanics are the same as any prompt or model evaluation. Build a dataset that serves as ground truth, run candidates against it, and score them. Use Chalk functions to define the dataset, create a task and scorers, and make sure every run is stored. For agent behavior you need the trajectory rather than the answer. Agent traces capture what the agent planned, which tools it called, and what those calls changed, so a run can be reconstructed and scored after the fact. What data access does an agentic ML loop need? An agentic ML loop needs access to the same data the production model uses, computed the same way, straight from the source. A loop confined to a single warehouse can only rediscover signal already in that warehouse. Agents that reach operational databases, streams, and APIs can find signals from the state of the world. In practice, hunting for a signal nobody has used means pulling from Postgres and Snowflake and Kafka and S3, plus an internal service and whatever vendor API somebody wired up last quarter. Often on the same afternoon. What the loop needs is one layer for context that assembles data from the source itself. What makes an agentic ML loop trustworthy? Many teams acknowledge the need and benefits of adopting agentic ML. But what decides whether you can trust its output is the data layer under it, and two properties in particular: that the agent sees what production saw, and that anything it proposes can be served without a rewrite. With Chalk's Context Engine teams define a feature once, in Python or SQL. That definition is what runs in training, in real-time serving, and in anything an agent reads. Values are computed from the source at the moment a decision needs them. Lineage and governance live in the same system as the context. Point-in-time correct training sets mean an agent backtesting a feature sees what production would have seen at that timestamp, rather than data that arrived later. One definition across training and serving means a feature the agent invents is servable by construction, so the loop cannot propose something that dies on the way to production. Computing from federated sources rather than a single warehouse is what lets the agent reach real-world signal without ETL jobs and data pipelines. Agents multiply whatever a team already has, which is why the data layer decides the outcome. An ML Autoresearch Agent can complete this loop against features computed the same way production computes them with Chalk's Context Engine. See what the loop looks like on your own data. Explore building an ML Autoresearch Agent with Chalk, or talk to an engineer about the model you would point it at first. Frequently asked questions What is agentic machine learning? Agentic machine learning is agents running the model development loop: proposing a hypothesis, implementing it, evaluating it against a metric, and keeping or discarding it based on the result. A human sets the objective and the agent optimizes toward it. It is the loop ML engineers have always run, executed at a rate a person cannot sustain. How is agentic machine learning different from AutoML? AutoML searches a space you defined in advance, across model families and hyperparameter ranges you chose. An agent can change the space itself: redefine a label, resample the data, reformulate a feature, or reframe the problem. AutoML optimizes within your hypothesis. An agent generates hypotheses. Does agentic machine learning replace ML engineers? No. The human keeps the objective and the approval. What changes is where an engineer's time goes: from writing and running experiments to setting the target metric and judging proposals. Engineers can review far more proposals than they can write, which is why the ceiling moves. What data access does an agentic machine learning loop need? The same data the production model uses, computed the same way, plus the sources that have not been centralized yet. A loop confined to one warehouse can only rediscover signal already in that warehouse. A loop that reaches operational databases, streams, and APIs can find signal nobody has used. Why do point-in-time correct training sets matter for agentic ML? Because an agent evaluating a feature has to see what production would have seen at that timestamp. If later data reaches the backtest, the agent scores the feature on information the model would not have had, accepts it, and the result does not hold in production. What should teams put in place before adopting agentic machine learning? A review cadence, an evaluation setup, and data access. Decide what level of change needs human approval, decide how you score both the model change and the agent's own behavior, and make sure the agent computes features from the same sources and definitions production uses. # Context Has a Timestamp: Real-Time Agent Context with Chalk source: https://chalk.ai/blog/context-has-a-timestamp Why real-time, windowed context beats static retrieval for agents — illustrated with a poker example that becomes a fraud detector. The most powerful input to an agent isn't the system prompt. It's your data. And not stale data. It's a snapshot of the world as it exists at the moment of request through numbers. Agent context is a feature value computed in real-time. Most agents access data through retrieval. The data is right. But it's static. All-time total, this month's average, today's count - audited, backfilled, and exact to the decimal. These numbers are true but the one that matters is the one calculated right now, using all the data available. A lifetime average computed last night is stale, and the same average computed at the decision isn't. I recently spoke at Temporal's "Agent Context is Everything" developer meet-up using poker as an example. Disclaimer: Chalk doesn't condone gambling, and Chalk has nothing to do with poker. That being said, poker strategy is about making decisions based on the changing signals you receive about other players, in real time. So, let's build an agent that plays poker intelligently. Poker Basics: A really quick intro so the rest of this makes sense. You never see an opponent's cards. Every bet is a guess about what they have. And there's a mathematically optimal baseline for that guess: a GTO (Game Theory Optimal) chart, which is a big lookup table that says: in this exact spot, fold this often, call this often, raise this often. It's a probability distribution. Most importantly, Nobody actually plays GTO. Not you, not the pro. The chart is a prior. The money is in measuring how far this person you are currently playing with has drifted from GTO, right now, to inform your bets, your decision to stay in, and your game strategy. A Hand Let's say I've been playing with the same people for 5 hours. That means I have 5 hours of intuition on these players. That's important for later. But right now, all that matters is the current hand. The board is K♥ 7♣ 2♦, rainbow. Three different suits, so no flush draw. My cards are A♥ K♦. Top pair, best kicker. Strong, but not a monster. My opponent bets small. The entire decision: do I call, raise, or fold? Layer 1: GTO Chart I look this spot up in my GTO chart. It says fold 0.08, call 0.62, raise 0.30. That means I'm supposed to call 62% of the time. Basically, the book says to call. That's the answer you get out of any agent whose context is static - a document. It's a lookup. But let's look at what layer 1 is missing. There's no opponent in it. We have 5 hours of data on how our opponent has played. With Chalk you can ensure those 5 hours of play affect our decision now. 4 Layers of Context The same opponent, four ways. Only the bottom row knows what's happening right now. Layer 2 is lifetime. Every hand this player has ever played in the demo scenario: 2,800 hands, VPIP 18.1%, 3-bet 3.1%. (VPIP is how often he voluntarily puts money in preflop. 3-bet is how often he re-raises.) That's a rock. Tight, patient, doesn't fight back. If those were the only numbers you had, you'd raise him off that small bet without thinking twice. He basically bets when he has it, and doesn't when he doesn't. Layer 3 is tonight's session. 300 hands, VPIP 19%. Same guy. The session confirms the lifetime number, which feels reassuring and tells you nothing you didn't already know. Layer 4 is trailing windows, anchored to the moment of the decision. In the last two hours: 120 hands, VPIP 24.2%, 3-bet 8.3%. Something's up. When we zoom in to the last thirty minutes: 30 hands, VPIP 40%, 3-bet 23.3%. His 3-bet rate went from 3.1% over his lifetime to 23.3% in the last half hour. More than seven times. That's not a rock. The "aggression factor" is 7x. That's a guy on tilt, right now. "On Tilt" is just poker speak for "the guy's angry and isn't thinking rationally." Before we look at what that implies, let's talk about how we got these numbers. The Code The fourth layer - the trailing windows - is one simple declaration in Chalk. \# src/features.py class PlayerSession hands_recent: Windowed[int] = windowed( "10m", "30m", "2h", expression=_.participations[ _.at >= _.chalk_window, _.at <= _.chalk_now, ].count(), default=0, ) vpip_hands_recent: Windowed[int] = windowed( "10m", "30m", "2h", expression=_.participations[ _.voluntarily_entered == True, _.at >= _.chalk_window, _.at <= _.chalk_now, ].count(), default=0, ) python 10m, 30m and 2h are declared inline, and Chalk computes all three on demand, straight off the raw rows. This is the entire way you capture a complicated materialized aggregation in Chalk, which is how you attribute recent actions to a player's overall strategy. There is no streaming job in this repo. No topic, no rollup table, no cron, no materialized view sitting somewhere with those numbers in it. It's simply a definition of how to compute a number. Our agent asks for that number right now and Chalk handles everything else, returning the most up to date number possible. In other words, a snapshot of the current state of the world in numbers. The agent \# cmp/advisor.py @chalkcompute.function( secrets=[ Secret.from_chalk_integration("pg"), ], image=Image.debian_slim(python_version="3.12").pip_install( [ "chalkpy>=2.130.5", "openai", "psycopg2-binary", ] ), ) def advise_hand( spot_prompt: str, spot_json: str, hand_id: str, chart_key: str, villain_player_id: str, player_session_id: str, now_iso: str, ) -> str: The image and the secrets are declared in place, on the function. The decorator is the deploy. There is no Dockerfile in this repo. No cluster, no service, no deployment manifest. Simply running the file ships the function, and the sandbox comes with its own identity. Plenty of platforms will host a sandbox for you. What I care about is how little of it I had to describe. The runtime credentials are named right there on the function, I registered them once with a Chalk secret set, and the platform injects them into the sandbox at run-time. And then the tool itself: \# illustrative - single-tool shape, not in this repo def chalk_query(inp: dict) -> str: ctx = chalk_client.query(input=inp["input"], output=inp["output"], now=now) return "\n".join(f"{a.field}: {a.value}" for a in ctx.data) The tool is literally a Chalk query. The model picks the feature list, and now= comes from the enclosing call, not from the model. That's the whole agent. A sandbox and a single tool - the ability to Chalk query. The windows are computed as of the instant of the hand being decided, not precalculated. In poker, freshness and latency matter. A number that doesn't take the last 5 hands into account simply doesn't matter. A number that's correct but arrives after the shot clock is worth the same as no number. Chalk enables real-time context to be served to agents. Outcome The GTO prior for this spot, against the distribution the agent returned. At the instant of the decision, the model had all four layers in front of it - the chart row, the lifetime profile, tonight's session, and the opponent's last thirty minutes. That last one was computed at the moment it was asked for, straight off the raw hand history. Nothing had written it down in advance. As we can see, with recent opponent history our decision changes. Our agent says we should raise 71% of the time. That's a 41% delta. Simply adding a real time context engine changes the distribution by 41%. That's an entirely different "best answer". Why Chalk One declaration, three live horizons. The windowed() call is the feature. It isn't the interface to a pipeline that computes the feature. There's no orchestration logic to maintain, and no data pipelines that can break. Federated, not materialized. The hand history lives in a database and it stays there. Chalk queries the source of truth at request time instead of precomputing the aggregates into a cache and then owning the problem of keeping that cache in sync. Low-latency serving. Serve the current state of the world within the time bound of a poker action. Enabling feature serving in single-digit milliseconds. Designed for high-throughput production workloads: 100,000 QPS with <5ms latency. Aggregate of an aggregate, no DAG. Counting a derived boolean that is itself a count, inside a time window, is the kind of thing that in most stacks turns into three jobs and an ordering problem. Chalk makes incredibly complicated computations simple. Swap the Nouns If we swap some nouns: Opponent becomes cardholder. Hand becomes transaction. GTO baseline becomes model expectation. And suddenly, we have an agent that detects changes in fraud patterns. Nothing else changes. Not the feature classes, not the windows, not the tool loop. A cardholder who's behaved one way across 2,800 transactions and a completely different way for the last thirty minutes is the same computation as a poker player on tilt. An agent with Chalk knows what just happened. Chalk enables real time context to be served in milliseconds to your agents. If you're looking for ways to assemble and serve agents and models fresh, real-time context, reach out anytime. # What is a Context Engine? | Chalk source: https://chalk.ai/blog/what-is-a-context-engine An agent only knows what's in its prompt. A context engine computes the values that describe a customer's situation right now and assembles them at inference time, so the agent stops guessing at facts the business already has. An agent that can't see current state doesn't fail loudly. It fails confidently and plausibly. Three examples that might sound familiar: - A support agent processes a refund for a customer with three opened disputes. It should have blocked it and triggered de-escalation. - A financial advisory agent quotes a portfolio balance that was correct an hour ago. Retrieval worked as designed. The index was just old. - A retention agent offers a win-back discount to a customer the risk system just froze. The agent scored on lifetime spend, while risk saw net chargebacks. Each of these agents was asked to make a call about a specific customer at a specific moment, and each was handed a pile of documents or a table. In every case the business already had what it needed to get the answer right. The building blocks of the right decision existed somewhere, but didn’t get to the model in time. All three are missing the same thing: a system that computes values and data relevant to customer state at the instant the request arrives, and assembles it in the prompt at inference time. That's a context engine. Great context makes prompts smaller and AI less wrong. Here’s how: Everything the agent knows about you is in the prompt Agents are powered by large language models. Ask any LLM about current events, and its frozen worldview quickly shows. Real-time accurate answers require a streaming view of the world, meaning something external must provide that information. An agent's current view of the world is strictly limited to its prompt at call time. Everything specific to this business, this account, this moment either arrives in the prompt or doesn't arrive at all. All action the agent takes is inference over a snapshot, the context in the prompt. This means most agent failures are assembly failures. Model quality is one variable. The other is how you build the world the agent sees on every reasoning loop: how close it is to what's true right now, and whether you can replay what the agent saw. Most of what gets loaded into context windows today are documents, because documents are what the current tooling makes easy to fetch. The values that describe what's happening right now are the harder half, and in durable systems, these decide the agent’s path forward. Why agents miss Three pathways put data in front of most models powering an agent. A vector index over documents. A memory store over history. And tool calls, which work differently from the first two: the agent reads its prompt, decides it needs something, calls an API, and the result gets appended to the context for the next turn. Tool calls are how an agent fetches context it wasn't given, one round trip at a time. None of the three paths were built for live data. Most retrieval infrastructure was designed to make written knowledge findable and it does that extremely well. It was never designed to compute a value from real-time data at the instant of a decision, so it doesn't. That gap shows up in three ways for fast moving companies: Business and world state. The agent sees what documents, databases, and semantic layers say, which is not the same as what's true at the moment it runs. A nightly load can't answer "what is this account's status this second." It also can’t know "is the payment corridor mentioned in this doc down right now" before it fails. Situational judgment. The agent gets raw records and is asked to infer. It never receives the derived values the business already uses to judge this kind of case, so it re-derives them badly or skips the step. Your risk team spent two years on a chargeback model. The agent is doing arithmetic on a transaction list. Workflow. Nothing tells the model which internal path to take. Escalate or resolve. Refund or withhold. Disclose or don't. With no deterministic signal, it improvises, and there's no record of how it decided. Dumping in more context is worse The instinct when an agent gets something wrong is to give it everything. Longer context windows make this feel free. It isn't, from both a cost and wasted outcomes lens: - Chroma tested 18 production models including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found accuracy falls as input grows. On LongMemEval every model did better on a focused ~300-token history than on the full ~113k-token log, and a single distractor was enough to pull it down. (source) - "Lost in the Middle" (Liu et al) found the shape of it. Models use what's at the start and end of a long context far better than what's buried in the middle, including models built for long context. (source) - Building legal agents, Harvey AI found that capping how many tokens a tool could return produced large gains across open and closed models, before any training at all. Their base model, with no way to know which documents mattered, bulk-read everything: 104 tool calls, 461k tokens, a 0.061 rubric score. Trained to be selective, 42 calls, 250k tokens, 0.803. (source) The reasoning models of agents aren’t short on context. They’re short on a few real-time values that would have told them what mattered most. Teams mostly know this. What blocks them is plumbing. An automation company described their invoice extraction to us as a set of silos: a vision model pulls line items with no access to that customer's purchase history, and a separate statistical layer catches implausible results afterward. When something looks wrong the invoice routes to a human, and the customer ends up touching the same document three times. What they wanted was to feed history into the extraction step so the model corrects itself. The real-time value was calculable in their systems. The model was never seeing that context. What engineered context looks like Take the support agent from the top. Before the prompt is assembled, the system resolves a dozen values for this customer: tenure, plan tier, open disputes in 30 days, chargeback risk, whether the jurisdiction is regulated, and failed payments in the last hour. That last one is specifically interesting, because it's the kind of value nothing else in the stack can produce. It isn't stored anywhere. It's an aggregation computed at inference time over a stream of payment events, some of which landed forty seconds ago. No document contains it. No nightly table has it. The warehouse won't know about it until tonight. The agent needs it now, because a customer whose card failed three times in the last hour is in a completely different situation than one whose card failed three times last spring. Those values don't inform the agent. They direct it. - Chargeback risk above threshold with open disputes: the refund tool is withheld and the case escalates. - Failed payments in the last hour with low tenure: pull the last three transaction attempts. - Regulated jurisdiction: suppress two tools and inject the disclosure. What reaches the model is a dozen values, several of them computed on the spot, plus two targeted retrievals: transcripts of how this exact situation resolved well before. The full account record and six months of tickets stay where they are. The principle underneath: the feature value is the control plane for context assembly. It's how the agent learns what kind of situation it's in before it starts reasoning about it. Where this sits next to your stack A fair question at this point is whether you already have this. A few things are adjacent. - Vector databases do semantic recall over documents. No computation over live systems. - Memory layers hold what was said. Conversation history goes stale the moment things change. - Knowledge graphs hold structure and relationships, as current as their last load. - Storage-first feature stores serve consistent precomputed values, with freshness set by whenever the pipeline last ran. - Semantic layers govern what a metric means. They map warehouse tables to business concepts so that revenue means one thing everywhere, and resolve queries against that definition. Vendors often considered for the context layer do a few different jobs. Glean runs enterprise search across workplace systems. LangChain gives you orchestration to wire retrieval and tools together. Neo4j holds the relationship graph. Mem0 and Zep hold conversation history. All of it is real, and most teams need some of it. What none of them do is compute a value from a live operational system at the moment of the decision. They can tell an agent everything that has been written about a customer and nothing about what that customer did four minutes ago. Retrieval still finds the policy document. The context engine decides whether this particular customer should be shown it. Semantic layers are a newer layer worth being specific about. When an analytics team works with AtScale, Dremio, or Cube: they’re building a governed definition layer over the warehouse, built to make reporting consistent. That answers "what does this term mean." A context engine answers "what is true and relevant right now." Different jobs for different challenges. Wire the two together and the semantic layer gets better. A semantic layer holds the definition of churn risk and the right way to reason about it, and today that logic runs against warehouse tables loaded last night. With a context engine, it reasons correctly and resolves on demand against live data instead. The governance doesn't change. The definition doesn't change. The number stops being twelve hours old. Isn't this just a feature store? If you came here looking for a feature store, you're in the right place. You probably need what a feature store was supposed to be. First-generation feature stores are databases of precomputed values. A pipeline runs, values land in a table, serving reads them back. That worked when the consumer was a model scoring a batch of rows on a schedule, and the freshness contract was whenever the pipeline last ran. An agent is a different consumer. It makes one decision, about one customer, at one moment, and it needs values no pipeline has computed yet. Failed payments in the last hour can't be precomputed. The answer at 3:04pm and the answer at 3:05pm are different computations over different events. A context engine is a feature store that computes instead of stores. Same feature definitions, same point-in-time correctness, same offline and online consistency. The difference is that the value resolves when the request arrives, rather than being read back from whenever the pipeline last ran. The feature store you built for your ML pipeline is the context engine your agents use. The primitives haven't changed None of the components for a context engine are new. It's the feature engine and model serving architecture that ML teams have run in production for years, pointed at a different consumer. - Agent and context are defined together in Python, so your AI and your models don't disagree about what a word means. - On-demand resolution from live sources. A value is current because it was computed when you asked. - Windowed aggregations, because "failed payments in the last hour" decides agent behavior and no document contains it. - Point-in-time correctness, which made training data honest and now makes evals replayable against the context production was served. - Millisecond latency, so the reasoning loop doesn't crawl. - Model serving and data lineage, so there's a record of what the agent saw and did. - Deployment inside your VPC, so sensitive data never crosses the open internet. Which values to build for agents Having the primitives doesn't tell you what to compute. Your user base has thousands of possible attributes. Which do you build out first? A useful filter: for each candidate value, name what agent path it decides. If a value doesn't change what the agent does, retrieves, or refuses, it's just decoration. Four classes are worth building, roughly in this order. Two are lookups. Two are calculations, and the calculations are where the work is. Entity state. The slow facts the agent keeps getting wrong. Tenure, plan tier, jurisdiction, account status. Cheap to define, and they resolve against the live source. @features class User: id: str tenure_days: int plan_tier: str python Windowed counts. A windowed aggregation runs over a stream of events at request time, bounded by a window that moves with the clock. A one-hour window at 3:04pm and the one-hour window at 3:05pm are different computations over different sets of events. Nothing stores the answer, which is why nothing can look it up. @features class User: failed_payments: Windowed[int] = windowed( "1h", "24h", "30d", expression=_.payments[_.status == "failed"].count(), materialization={"bucket_duration": "10m"}, ) Derived judgments. The scores and tiers your business already computes, usually inside a model or a SQL view no agent can reach. Define them once so the agent and the risk system can't disagree. @online def get_chargeback_risk( failed_1h: User.failed_payments["1h"], disputes_30d: User.open_disputes["30d"], tenure: User.tenure_days, ) -> User.chargeback_risk: return risk_model.predict(failed_1h, disputes_30d, tenure) Routing flags. The booleans that gate tools. This is where the control plane stops being a metaphor. @online def get_refund_eligible( risk: User.chargeback_risk, disputes: User.open_disputes_30d, ) -> User.refund_eligible: return risk < 0.4 and disputes == 0 The agent doesn't need to reason about whether to offer a refund here. It reads a boolean computed the same way the refund system computes it. These last two classes also decide what to retrieve Gating tools is the obvious use. The less obvious one is that the same values should be choosing which documents the agent sees. Say `refund_eligible` resolves false and the case is heading toward a denial. What helps is the ten past denials in this product line that customers accepted without escalating, plus the two policy sections that apply to this jurisdiction and plan tier. The full policy library does nothing. Most stacks retrieve first and reason later, which means the model gets a semantically similar pile and has to work out which half applies. Resolve the values first and retrieval runs over a shelf instead of a warehouse. Same index, far better inputs, and the agent stops spending its reasoning budget figuring out which situation it's in. Tying it all together Let's revisit the three agents from the beginning. The support agent never saw the fraud model. It didn't need to. The fraud model's output is a derived judgment, `chargeback_risk`, and the agent reads it through a routing flag: `refund_eligible` resolves false, the refund tool never enters the tool list, and the case escalates. Two of the four classes, chained. The agent didn't get better at judgment. It stopped being asked to have any. The advisor quoted an hour-old balance because the balance was entity state served from an index. Resolve that same entity state against the system that owns it and it gets computed when the question is asked. There's no window during which the answer is quietly wrong. The agent that called a blocked customer high value needed one definition instead of two. `value_tier` is a derived judgment computed once and read by both the agent and the risk system, so the two can't contradict each other in front of the customer. None of these are model problems. Swap in a better model and it gets all three wrong the same way, because none of them were errors of reasoning. The information needed never made it into the prompt. This is already running These systems aren't theoretical. Mercury, Whatnot, Socure and more run risk decisions, recommendations, embeddings, and agents in sandboxes on context engines today. Typical time to compute a value is 5 milliseconds. The question isn’t if your AI stack has the time, it’s whether you can build reliably without it. So, do you need a Context Engine? Not always. If your agent works over static documents and a nightly refresh is just as accurate, you don't need this. Check below. You know you need a context engine if any of these describe your agent: - It acts rather than answers. Issuing a refund, approving a limit, releasing a shipment, routing a case. An action taken on stale information is a decision the business has to unwind. - The answer changes within a session. A balance, an inventory count, a fraud score, a rate. If a value can move between the user's first message and their third, an index can't carry it. - Recency changes the meaning. Three failed payments this hour and three failed payments last spring are the same number describing two different customers. - Another system already has an opinion. A risk model, an eligibility service, a pricing engine. If the agent computes its own version, the two will contradict each other in front of a customer eventually. - Someone will ask what it saw. Regulated decisions, disputes, audits. Reconstructing why an agent did something requires knowing what was in the prompt at the time. None of these apply to an agent that summarizes analysis. Most of them apply to anything touching a customer account. AI is in a paradigm shift. We're handing agents real decisions while keeping them ignorant of the situation they're deciding about. The systems that fix this already exist. They were built for fraud and risk teams who couldn't afford a stale answer, and they deliver context consistently and point-in-time correct under load. Pointing them at agents is the shortest path to an agent that knows what's going on. # What is a Feature Store? source: https://chalk.ai/blog/what-is-a-feature-store A feature store is more than a tool – it is a secret weapon that empowers teams to innovate faster, collaborate better, and deliver production grade machine learning. A feature store is a centralized system that manages and serves machine learning features, the transformed data that models use to make predictions. It ensures features are defined once and can be consistently reused across training and production, keeping models accurate and reliable. When ML teams first come to Chalk, they rarely describe the problem as a feature store problem. They describe a bug. A model that scores differently than it did in testing. A discrepancy between what a batch job computed overnight and what the API is returning right now. A debugging session that ends with two engineers on two different teams staring at two different implementations of what was supposed to be the same feature. The pattern shows up often enough that it has a name: training-serving skew. It rarely looks obvious at first, because each version of the feature is correct in isolation. They just drifted apart quietly, one change at a time, until the outputs stopped matching. This is the problem a feature store exists to solve. By acting as the single source of truth for features, defined once and reused across training and production, a feature store closes that gap and lets engineering teams move faster, collaborate more effectively, and scale models with confidence. What is a feature in machine learning? In machine learning, a feature is any measurable piece of data that a model uses as an input to make predictions. Features aren't the raw data itself, they're derived from it. Raw data is the source material; features are what you get after transforming, aggregating, and engineering that data into a form a model can learn from. For example: - A bank predicting fraudulent transactions might engineer features like "transaction amount relative to the customer's 30-day average" or "number of foreign transactions in the past week," both derived from raw transaction logs. - An e-commerce team building a recommendation model might create features like "number of purchases in the last 30 days" or "average time between sessions." The process of creating these features from raw data is called feature engineering. It's one of the most time-intensive parts of building ML models, and managing those features consistently across training and production is exactly what a feature store is designed to do. Why feature stores matter for ML teams Machine learning teams often discover that the hardest part of building models isn't the model itself, it's managing the data that powers it. Without a feature store, organizations face recurring problems that slow down delivery, increase costs, and make models less reliable. Common challenges include: - Inconsistent training vs production features: Features are often reimplemented when models move from notebooks into production, which leads to drift and mismatched results between training and inference. - Duplicate work across teams: Different teams frequently rebuild the same features in parallel, wasting time and slowing experimentation velocity. - Inefficient recomputation: Complex transformations get recomputed repeatedly instead of being reused, driving up infrastructure costs and delaying iteration cycles. - No single source of truth: Without a central system, feature definitions live in scattered scripts and pipelines, making governance and compliance nearly impossible. - Hard-to-debug lineage: When something goes wrong, teams struggle to trace how a feature was built, what data sources it depended on, and where it diverged. This erodes trust in models once they're deployed. Feature stores were built to solve these problems. By centralizing feature definitions and making them reusable across training and production, they ensure consistency, reduce wasted effort, and give ML teams the confidence to scale models into production. Core benefits of a feature store Adopting a feature store solves not only common pain points, but it unlocks new capabilities that make MLOps teams faster, more consistent, and easier to trust. - Consistency: Feature stores ensure training and production use the exact same feature definitions. This online/offline sync eliminates drift, so predictions in production match the results you saw during model development. - Real-time serving: Fresh features can be delivered in milliseconds, powering critical applications like fraud detection, personalization, and recommendations. Instead of relying on stale batch data, teams can make decisions with the most up-to-date signals available. - Faster experimentation: By centralizing and reusing features, teams can build new models without starting from scratch. This accelerates iteration cycles, reduces duplicate work, and helps organizations scale experimentation across multiple teams and projects. - Feature lineage & governance: A feature store tracks how features were created, their dependencies, and how they've evolved over time. This lineage makes debugging easier, supports compliance requirements, and builds confidence in the reliability of production models. Common use cases for feature stores Feature stores are increasingly seen as core infrastructure because they make it easier to deliver reliable, low-latency data to models. Here are some of the most common applications in real-world ML systems: Fraud detection Real-time fraud systems rely on streaming features like transaction history, device fingerprints, and geolocation signals. With a feature store, these inputs are served in milliseconds, helping financial institutions flag suspicious activity before it reaches the customer. For a detailed case study on how a feature store can be used to build a fraud detection system, refer to Feature Store at Work: A Tutorial on Fraud and Risk. Recommendations Recommendation engines depend on up-to-date behavioral data (what a user has watched, clicked, or purchased recently). A feature store ensures those signals are fresh and consistent, powering more accurate and relevant recommendations. Personalization From media streaming to e-commerce, personalization requires fast access to a user's latest activity. Feature stores can stream events like likes, shares, or browsing history at low latency, enabling models to respond in real time. Compliance In regulated industries like finance and healthcare, reproducibility is critical. Feature stores support "time travel," making it possible to reconstruct what a feature looked like at any point in time, essential for audits, debugging, and compliance reporting. These use cases show why feature stores are now standard infrastructure for modern ML teams. Feature store vs data warehouse vs vector database When exploring ML infrastructure, it's easy to confuse feature stores with other systems like data warehouses or vector databases. While they sometimes overlap, each serves a distinct purpose. - Data Warehouse: Data warehouses excel at batch analytics, business reporting, and aggregating historical data. They aren't designed for low-latency inference or temporal consistency, which are essential for ML models in production. - Vector Database: Vector databases are optimized for similarity search and embeddings, great for use cases like semantic search or recommendation retrieval. However, they don't provide feature engineering, consistency guarantees, or feature lineage tracking. - Feature Store: Feature stores are purpose-built for ML workflows. They handle feature computation, ensure online/offline consistency, enable feature reuse, and provide lineage so teams can debug, audit, and govern features confidently. Data Warehouse Vector Database Feature Store System Batch analytics, business intelligence Embedding storage, similarity search Feature computation, serving, consistency, lineage Strengths Not built for real-time inference or time travel Doesn't handle feature engineering or lineage Purpose-built for bridging training and production Limitations for ML Workflows Each system has its place, but only feature stores are designed to reliably bridge the gap between training and production. Feature store architecture explained Feature stores are defined by four core components in the flow of data: - Sourcing: connecting to raw data sources, - Transforming: loading and running transformations on data, - Storing: persisting transformed features, and - Serving: providing access to transformed data. There are two additional (though equally important) components that extend naturally from the core components and respond to the challenges that feature stores address: - Monitoring: detecting feature drift, shifts in latency and storage usage, and data freshness. - Experimenting: iterating on feature and resolver definitions during model training. These two components aren't definitional, but they are essential parts of making transformed features production-grade. Sourcing Feature stores are built to be data source agnostic. They implement connectors that allow data to flow in from your company's upstream data sources, making them easy to integrate into any existing data architecture. Common categories of data sources include real-time, streaming, and batch data sources from which the feature store can load feature data and all the input data required to generate feature data. Transforming Feature stores compute and persist transformed data. Users write most of their custom logic and code for the transformation step of a feature store. The goal of transformation is to specify the desired structure of your features and the logic with which to compute feature values. These specifications comprise the registry for the feature store, which serves as a schema for all current and historical feature definitions. The registry should contain definitions for the input data sources or features and specific computations required to determine a feature value. Once the registry has been defined and deployed, feature stores further optimize performance by performing these data transformations responsively. Rather than running all the computations specified in the registry proactively on all possible upstream data, a well-implemented and efficient feature store runs transformations on demand to compute data that has not already previously been computed and persisted in its stores. When a feature store receives a query, it determines which feature values are already computed and stored. For the values that need to be freshly computed, it determines the registry definitions that map to each feature value, runs those specified transformations, and then returns and persists the requested feature data. Storing Machine learning teams require two forms of data access: real-time, to serve modeled or computed results to customers, and retroactive, to generate data for model training, to monitor how features are changing over time, or to run analytic queries. These are incredibly different access patterns, but their outputs need to be consistent. The data used to train a model must look like the data used to make predictions. Combining them under one interface guarantees this consistency and improves efficiency. Feature stores like Chalk accomplish this through two abstractions: an online store and an offline store. The online store is a key value store responsible for low latency serving of features in real-time. It remembers the latest computed values of your features based on their primary keys. The offline store is responsible for remembering every feature you've computed. It can also be queried, for instance, to access historical data. Feature registry The feature registry is the centralized catalog of all feature definitions within a feature store. It acts as the single source of truth for every feature in your organization, storing not just the feature values themselves, but the metadata that describes them: how a feature is computed, what data source it draws from, how frequently it updates, who owns it, and which models depend on it. The registry is what enables feature discovery and reuse. Before building a new feature, a data scientist can search the registry to find out whether something similar already exists, and if it does, they can reuse it directly rather than rebuilding from scratch. This is one of the primary ways feature stores reduce duplicated engineering effort across teams. In Chalk, the registry is defined through your feature class definitions and resolver logic. When you deploy a feature, its definition, dependencies, and lineage are automatically recorded and made queryable, giving teams full visibility into what's running in production and how each feature was derived. Serving Feature stores typically provide an interface for requesting data from either the offline or online store. At a high level, access to a feature store is divided into online queries and offline queries. The goal of online queries is to return the latest value for a feature as quickly as possible, either by returning the values from a cache or by running the necessary resolvers to calculate the requested features. The goal of offline queries is to either: 1. Retrieve historical data that has already been calculated, 1. Warm the online store cache (through an offline-online ETL process), or 1. Run batch jobs to generate features that don't need to be served through the online store. There are multiple ways to run offline and online queries, but some of the most common methods include: - Through clients (implemented in different programming languages), - By making HTTP requests to a REST API, - Through scheduled orchestration of queries, which will run queries on a schedule similar to scheduled ETL jobs. Monitoring Monitoring is a critical, yet less rigorously defined, part of feature stores. Knowing how your features and resolvers are behaving (or misbehaving) can surface previously invisible or inaccessible information, revealing subtle bugs early. Because feature stores centralize feature computation, they can provide a comprehensive view into your features, including: the latency of your queries, the relationships between your features, the consistency of a feature's distribution over time, and the number of times a particular resolver is being run. A good feature store allows you to define very granular metrics and connect them to your existing monitoring and alerting systems. For instance, if you have a feature defining whether a transaction is likely fraudulent, you could configure a monitor that alerts an external service if the percentage of queries marking transactions as fraudulent is suspiciously high or low. This allows instant visibility into critical decisions made by your ML system that otherwise might be hard to detect. Experimenting A well-implemented feature store lets you experiment and collaborate on features and pipelines. Typically, clients create multiple environments for development and testing for their feature stores. This allows feature transformation code to be tested and evaluated before it is embedded in production pipelines. In the same vein of testing in isolated environments, feature stores can also enable engineering best practices for collaboration through isolated deployments within an environment, similar to version control, allowing for concurrent iteration. What is training-serving skew? Training-serving skew happens when the feature values a model sees during training don't match the feature values it sees in production, even though both were supposed to represent the same thing. It usually starts small. A data scientist builds a feature in a notebook, using Python and Pandas, to test a new model idea. The model performs well. To ship it, an engineering team reimplements that same feature logic in the production serving path, often in a different language, on a different schedule, against a live data source instead of a static training set. The two implementations are supposed to compute the same value. For a while, they do. Then something changes upstream: a data source adds a new field, a business rule shifts, an edge case in the raw data shows up that the original implementation never accounted for. One version gets updated. The other doesn't. The two features quietly start returning different values for the same input, and nothing in the system flags it, because both versions still run without errors. They just no longer agree. This is difficult to catch because both versions look correct in isolation. Teams usually notice it downstream, as a model that performs worse in production than it did in testing, or a decision that seems inconsistent with what the data should support. By the time it's traced back to two divergent implementations of the same feature, it can take days of debugging across two teams to confirm what went wrong. A feature store addresses this directly by making the feature definition itself the single source of truth. The same code that computes a feature for training also computes it for serving, so there's no second implementation to drift out of sync in the first place. Feature store solutions and tools Feature store tooling generally falls into three categories, each suited to different team needs. Open source feature stores. Feast is the most widely adopted open source option, providing a schema and orchestration layer that you connect to your own online and offline stores. Open source options offer flexibility and no vendor lock-in, but usually require a team to own infrastructure, backfills, and scaling themselves. Managed cloud feature stores. Cloud providers offer feature stores that plug directly into their broader ML platforms, such as Amazon SageMaker Feature Store and Google Cloud Vertex AI Feature Store. These are a natural fit if a team is already standardized on one cloud's ML stack, though they can be harder to adopt if data lives across multiple platforms. Compute-first feature platforms. A newer category, including Chalk, treats feature computation as a first-class part of the system rather than something that happens upstream in a separate ETL pipeline. Instead of only storing and serving precomputed values, these platforms compute features on demand, at query time, from live data sources. This closes the freshness gap that storage-first feature stores run into and reduces the operational burden of maintaining separate pipelines for training and serving. The right choice depends on how fresh features need to be, how much infrastructure a team wants to own, and how many different systems feature logic needs to stay consistent across. Do you need a feature store? Not every team needs a feature store right away. For some, existing pipelines and data platforms may be enough. But as machine learning projects grow in scope and complexity, the need for a dedicated system to manage features becomes clear. When you might not need one (yet) If your team is only running a handful of offline models and feature definitions live comfortably within existing ETL pipelines, a feature store may be overkill. Similarly, if real-time serving isn't critical to your use case, you can often get by with a simpler setup. The tipping point rarely shows up as a checklist It shows up as an incident. A model has been running fine for months. Then something changes upstream, a data source, a schema, a business rule, and the batch feature and the real-time feature quietly stop agreeing. By the time anyone notices, the mismatch has already been feeding decisions for weeks. Teams that reach this point don't usually say "we need a feature store." They say "we need to find every place this feature gets computed and make sure they're all the same," and realize partway through that fixing it properly means building the kind of system a feature store already is. Signs you've reached that point: - You need consistency between training and production features to prevent drift. - Your team spends significant time rewriting Python code for production environments. - You require low-latency, real-time inference for fraud detection, recommendations, or personalization. - Multiple teams are rebuilding the same features, slowing iteration velocity. - You need feature lineage and time travel for compliance, debugging, or audit requirements. Eliminates training-serving skew, same feature logic in training and production Feature reuse reduces duplicated work across teams Real-time serving at low latency for fraud, recommendations, personalization Centralized lineage and governance for compliance and debugging Accelerates experimentation by giving teams a library of production-ready features Advantages Initial setup requires engineering investment and organizational buy-in May be overkill for small teams running only a handful of offline models Requires ongoing maintenance as data sources and feature requirements evolve Integration with existing pipelines can require significant refactoring Performance bottlenecks if not properly optimized at scale Disadvantages Feature stores were designed to solve these problems, giving ML teams a single source of truth for features and a reliable bridge between prototyping in notebooks and production deployment. Feature store FAQs A few things teams ask once they start weighing their options. Is a feature store the same as a feature engine? Not quite. A feature store serves stored, precomputed values and depends on ETL to keep them current. A feature engine computes feature logic on demand from live data, so features stay fresh and the training and serving paths stay consistent without a separate pipeline. Should you build a feature store or buy one? It comes down to how much infrastructure you want to own. Building in-house gives you full control and can be enough when your feature needs are narrow and stable, but you take on the pipelines, backfills, and scaling. Most teams move to a managed platform once several models share the same feature logic and real-time serving is in play. How is a feature store different from an MLOps platform? A feature store owns one layer: the data and features your models consume. An MLOps platform is wider, handling model training, deployment, monitoring, and retraining across the lifecycle. The two fit together, with the feature store supplying consistent inputs to the workflows your MLOps tooling runs. Conclusion Feature stores have become critical to machine learning platforms at companies of all stages, in all industries, and all sizes of ML and data teams. If your team is looking to ship production-grade machine learning confidently and quickly, a feature store is table stakes. Chalk is data source agnostic, supports client libraries across multiple programming languages, and keeps deploys and queries fast enough for real iteration during development. See what a best-in-class feature store feels like for your own team. Book a demo. # What Snowflake Summit Was Really About source: https://chalk.ai/blog/snowflake-summit-agents-compute What we heard at Snowflake Summit, and what we built to answer it. We spent last week at Snowflake Summit. The question we heard more than any other: how do we reliably trust agents in production? Four years ago, we bet that real-time data infrastructure would become as essential to AI as databases are to applications. That bet started with traditional ML. Now it runs through the question every enterprise AI team is wrestling with — getting dynamic agents into production safely. Chalk Compute is our answer: an enterprise runtime for AI agents, model inference, and other workloads. It runs each workload in a secure sandbox inside your own cloud environment, so your code and data never leave your VPC. Because Compute is built on Chalk's Context Engine, it does what general-purpose runtimes can't. It serves your agent historical production data, making production-grade agent evaluation possible before anything ships. When an agent runs, it makes a series of calls and queries. Every one of those steps is anchored to the same moment in time. The agent can't accidentally read data that doesn't yet exist. That's what makes real backtesting possible. What We Saw on the Show Floor Two things from the show floor are worth sharing. AJ Balance, Grindr's CPO, joined me on stage to walk through how they built an AI-first consumer company at scale: 15 million monthly active users, 65 engineers, and $440 million in revenue. His throughline was direct: the infrastructure question is the product question. Grindr became a globally recognized app by pairing the right features with the right architecture. AJ said it directly: Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project. What would have taken months took weeks. Separately, my co-founder Elliot Marx demoed Chalk Compute to a standing-room-only crowd, walking through what it looks like to lock an agent to a point in time and run it against real historical data. For engineers who have spent years stitching together synthetic test environments and shipping with their fingers crossed, Elliot captured their attention. Bridging The Temporal Gap Teams we work with are blocked on agent evaluation. Before any autonomous system touches a production workflow, one question has to be answered: how would this agent have behaved against the data it would actually have seen? For most teams, the honest answer is that they do not know. They build synthetic test environments, make educated guesses, and ship, hoping it works. When something goes wrong, reconstructing what the agent saw and why it decided what it did is either painful or impossible. Chalk Compute solves this by letting you lock an agent to any point in time and run it against the data your production system would have actually served at that moment. Not a synthetic approximation. The real thing, inside your own cloud. The clearest version of this problem is from Grindr. Their requirements were absolute: full data residency, no third-party routing of user data, and the ability to ship agents without triggering a manual security review every time something changed. That meant they needed to test agents against point-in-time production data. Chalk Compute runs entirely inside Grindr's VPC. Their trust-and-safety systems use agents to run in real time at scale. That’s just one reason they’ve become a leader in consumer AI. The teams that will win with agents are the ones who treat the data layer as load-bearing from the start. They learned this lesson with ML models. If you're building agents you need to be able to trust and evaluate in production, we'd love to talk. Request a demo. # Announcing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud source: https://chalk.ai/blog/chalk-compute Run agents against the exact data your production system would have served, in your cloud. Building agents is easy. Trusting them in production is not. The real bottleneck isn't raw container execution speed, it's the outer loop: evaluating your agent against real historical context, fixing what broke, and redeploying with confidence. Today we're launching Chalk Compute to close that loop. Chalk Compute is an enterprise-grade agent runtime that deploys sandboxes directly inside your private cloud. Tightly integrated with the Chalk Context Engine, it makes it simple to run temporally consistent agent evals on your historical production data. A single “knowledge cut-off” parameter initializes your sandbox and routes every tool flow through the Chalk MCP gateway, locked to that point in time. As far as the agent knows, it's running in the past. The Chalk Compute architecture is already hardened at production scale. We have spent years building production ML infrastructure for customers like MoneyLion, Whatnot, and Grindr. Grindr runs Chalk Compute directly in their cloud to orchestrate trust-and-safety workloads that protect more than 15 million users, and to run engineering pipelines where agents now handle 80% of internal code commits. The industry is optimizing the wrong loop. The broader AI ecosystem is locked in an infrastructure sideshow, racing to optimize the agent’s inner loop: making containers spin up marginally faster to read, decide, and act. But for teams deploying real enterprise applications, the bottleneck isn't raw container execution speed. It is the developer's outer loop: evaluating an agent against real historical context, fixing prompts or tools, and redeploying with confidence that agents will behave as expected. General-purpose runtimes force you to fabricate brittle synthetic test environments to guess at agent behavior. Because Chalk Compute sits natively on top of our real-time, temporally consistent context engine, it can deliver the counterfactual: how the agent would have behaved against the world it would have faced. There’s a well-worn claim: all production application problems are database problems. Agents haven't changed this. Context engineering is a new name for feature engineering, which is a new name for data engineering. The name changes, the job doesn’t: identify the signals that drive decisions, get them to the model instantly, and ensure they mean the same thing at training time that they do at inference time. Chalk has been solving that problem for years before "agent" entered the vocabulary. Chalk Compute is the execution surface. The Chalk Context Engine is the substance. The Chalk Context Engine: A Feature Store with a Memory Most feature stores only serve the present. They capture the current state of a database or the latest log line, leaving developers to patch together historical data when testing or auditing applications manually. The Chalk Context Engine serves the past as easily as the present. Every value it serves, whether it’s a customer’s transaction history, a fraud signal, or the output of a model call, is versioned by time. The Context Engine lets you ask, “What was this user’s score on March 14 at 3:42 PM?” and instantly get the same answer your production system would have served at that precise moment. This is what makes time-traveling agents possible. And it’s fast. To make this data useful to enterprise-grade agents and models, delivery must be instantaneous. Chalk’s native Python resolvers connect your existing production data sources—including Snowflake, Databricks, BigQuery, Postgres, Kafka, internal APIs, and external SaaS vendors—to serve features with sub-5ms latency. The Chalk Context Engine eliminates the need to build, scale, and maintain complex RAG retrieval pipelines or undergo massive data migrations. Your agents query live production data the same way your ML models do, against a temporally consistent view of your business state. We operate on a simple principle: your production data is your best training and evaluation environment.[1] True Backtesting with Temporally Consistent Agent Trajectories It’s important to understand how this differs from a traditional "as-of" query in a data warehouse. Those can give you point-in-time lookups on historical data. But they can’t control the agent's behavior. When an autonomous agent executes a loop, it ultimately takes actions: querying internal services, running sandboxed code, or hitting external APIs. This series of sequential actions forms what we call an “agent trajectory.” Every tool and data call your agent makes passes through the Chalk MCP Gateway. When a developer initializes a Chalk Compute sandbox with a knowledge cutoff, the MCP Gateway acts as an inline security and data governor. It doesn’t just fetch old database records; it propagates time bounds to all downstream tools and APIs so that they respond as they would have at that specified moment in history. An agent executing a simulation or a replay test can never accidentally peek into the future. By combining point-in-time consistency with active runtime enforcement, Chalk provides the only platform capable of executing truly time-locked, temporally consistent agent trajectories. Chalk Compute: Closing the Outer Loop at Enterprise Scale Chalk Compute deploys entirely within your private cloud. Your agents, your context queries, and your tool calls route exclusively through your own VPC infrastructure, never touching any third-party servers. The platform automates the heavy lifting: container builds, image caching, autoscaling, and secret injection. Every sandbox runs under gVisor, a user-space kernel that intercepts syscalls before they reach the host. Outbound egress is locked to a hostname or CIDR allowlist you control. And it's fast: it scales to 10,000 isolated containers in under 10 seconds to handle massive concurrent agent workloads. The platform delivers developer ergonomics without sacrificing enterprise security boundaries. Defining a secure agent sandbox and binding its entire downstream trajectory to a specific point-in-time requires nothing more than a little Python: from chalkcompute import Sandbox, Image from chalk import ChalkClient from agents import Agent, Runner # The sandbox sees the world as it was. sandbox = Sandbox(knowledge_cutoff="2025-09-15T14:30:00Z") @sandbox.function(image=Image.debian_slim().pip_install(["openai"])) def investigate_refund(order_id: str) -> str: # ChalkClient queries return the data as it existed at the cutoff. risk = ChalkClient().query( input={"order.id": order_id}, output=["order.refund_risk_score", "order.customer_prior_refunds"], ) # External calls are replayed from historical recordings, governed by policy Runner.run(Agent(tools=[CreditReportTool(), DatasourceQueryTool()]), "...") ... python With this architecture, a simple Python decorator replaces hours of infrastructure overhead. You no longer need to write custom Dockerfiles, manage complex Kubernetes manifests, or handle cluster administration. You can swap an LLM framework, modify a tool payload, or adjust a prompt in plain Python without ever touching a deployment configuration file.[2] Context Isn’t Simply a Compute Bolt-On General-purpose container runtimes like Modal, E2B, or Daytona can spin up sandboxes in the cloud. But none of them can rewind the world your agents read from. They serve the present, or a snapshot you’ve manually assembled. There is no mechanism to lock an agent’s entire tool trajectory—every downstream API call, every LLM context window—to a consistent past moment. That’s an architectural gap that can’t be patched onto a generic runtime. Attempting to recreate time-traveling agents on a generic runtime requires manually standing up an enterprise feature store, building a custom real-time retrieval layer, and engineering a stateful policy router to enforce knowledge cutoffs. We have spent four years hardening that exact real-time data delivery engine for the world’s most demanding production ML environments. Data delivery is the hardest problem in enterprise AI. It is the one thing you cannot simply bolt onto a compute runtime after the fact. Together, Chalk Compute and the Chalk Context Engine form the first unified AI data platform built for speed, hosted in your cloud, and designed to run autonomous agents against real-time and historical point-in-time data with absolute data sovereignty. Grindr: What You Can Ship When the Data Doesn’t Move When Grindr, already a Chalk Context Engine customer, started building agentic workloads, they needed a compute runtime to accelerate their developer outer loop. Their use cases span trust-and-safety agents, matchmaking agents turning multimodal user behavior into real-time recommendations, and internal engineering agents that now route 80% of the team's code commits. Handling some of the most sensitive PII in consumer technology means Grindr’s infrastructure constraints are absolute. Any agent platform that routes user payloads through a vendor's third-party managed infrastructure is an immediate non-starter, regardless of the security guarantees offered. They needed a platform that could execute untrusted code, enforce strict data governance, and maintain complete end-to-end visibility without a single byte of data ever leaving their VPC. Residency isn't the only hurdle at Grindr's scale. InfoSec review velocity is a major concern. Most agent platforms force a manual safety review for every new agent configuration or prompt change. This bottleneck scales linearly with deployments. Chalk Compute shifts the unit of security from the individual agent to the platform: permissions, audit logging, and MCP routing all live at the gateway level, so InfoSec clears Chalk once inside the VPC, allowing developers to ship thousands of distinct agents without new review cycles. Grindr Chief Product Officer AJ Balance describes the operational impact: "Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project. What would have taken months took weeks." Run Your First Time-Traveling Agent As the AI ecosystem matures from conversational chatbots to autonomous agents that actively manipulate business state, the critical architectural constraint has migrated from raw compute capacity back to data delivery. The teams that will define what agents can do in production are the ones who treat the data layer as load-bearing from the start. By unifying a real-time, temporally consistent context layer with secure, in-VPC agent sandboxes, Chalk Compute allows engineering teams to stop fabricating brittle synthetic test environments and start shipping with absolute historical confidence. If the architecture holds for Grindr's data, it'll hold for yours. If your eval set doesn't include the exact data your agent would have actually seen at inference time, it's time to fix your outer loop. - Enterprise Availability: Chalk Compute is available today for customers running on AWS (EKS) and GCP (GKE), with support for Azure (AKS) coming soon. - Open Source: This summer, we'll be open-sourcing the Chalk Compute runtime so engineering teams can inspect, extend, and run it themselves. Learn more about Compute and how to get started. 1. Because the Chalk Context Engine serves the same production signals to your training runs that your agents rely on at inference time, your proprietary data becomes your most powerful RL environment. More on this soon. 1. If you need the model call to stay inside your VPC too, simply swap openai:gpt-4o for a self-hosted model running on Chalk Compute (Qwen, Llama, whatever you're running). # Welcome Home, West Coast source: https://chalk.ai/blog/sf-office-warming-party Chalk celebrated the opening of its new San Francisco headquarters with 100+ engineers and builders from the Bay Area AI and ML infrastructure community. Three co-founders in one room. That's where Chalk started. Last Tuesday, we hosted 100+ people in our new San Francisco headquarters. Engineers, builders, and members of the Bay Area AI and ML infrastructure community came through to see the new space, meet the team, and share a drink (or three). VP of Marketing Alexandra Kane kicked things off with welcoming remarks, setting the tone for a night of smart people, good conversations, and genuine excitement about what's being built in AI infrastructure right now. Guests arrived to a spread of hors d'oeuvres and high-top tables stacked with Chalk swag, the kind of display that made it hard to walk past without grabbing something. From there, the night took on a life of its own. People wrote messages on our welcome board, which will be framed and hung on the wall as a permanent piece of Chalk's history, and settled into games of chess, poker, and mahjong that got surprisingly competitive as the evening went on. For a lot of attendees, it was also their first deep dive into who Chalk is, and we loved that. If you're an engineer or working in AI and you've been curious about us, this is what we're here for: see the space, meet the people building it, and hear firsthand about the problems we're solving. Because we're hiring. Across the company, across disciplines. If you're excited about the AI infrastructure space and want to work somewhere that's still early enough to matter, we want to talk to you. April also marks the opening of our Los Angeles office, joining San Francisco and New York. The West Coast is officially covered, and we're not slowing down: look out for us in Europe and Asia soon. Want to meet us in person? We'll be at events across SF, LA, and New York in the months ahead. Follow us to stay in the loop, and if you're interested in joining the team, check out our open roles below. View open roles at Chalk. # Chalk's Spring Fling: Where to Find Us This Season source: https://chalk.ai/blog/spring-fling-schedule Chalk's Spring Fling is its most ambitious events season yet. From an SF office warming party to Bronze sponsorship at AI Council, a first-time presence at Snowflake Summit (booth #2611, June 1–4), and something in the works at Databricks Data & AI Summit, Chalk is showing up everywhere AI and data infrastructure is being built. Spring is usually when everyone exhales. Not us. This quarter, Chalk is showing up everywhere that matters: conferences, around the streets of San Francisco, hackathons, dinners, and a few stages we've been working toward for a while. We're calling it Spring Fling because the team will not let it go. (I tried to land on something more grown-up. I lost.) The most ambitious stretch of events we've ever run, packed into eight weeks. Why now? Because the category is moving faster than at any point since we started this company. Our team has doubled in the last year. We just had our best pipeline quarter ever. And the buyers we care about, the people building production AI and ML at scale, all show up in person this season. So we will too. It all kicks off today. Welcome Home, West Coast Tonight, we're hosting our official San Francisco HQ Office Warming Party. The Bay Area tech and ML infrastructure community is invited. Come for the food, stay for the people, get a look at the new place. If you're building in AI or data infra and you're local, come find us. April also marks the opening of our new Los Angeles office, Chalk's third location, joining San Francisco and New York. Three cities feels a little surreal for a company that started with three co-founders crammed into one room. But here we are. Taking the Stage In May, we step onto the conference stage as a Bronze Sponsor of AI Council (May 12–14), a premier gathering for the AI and data leaders defining where this category goes next. It's an audience we've been wanting to be in front of for a while. We're proud to be there. Same week, keep your eyes open around San Francisco. You'll be seeing us in a few unexpected places. Keep an eye on in May and June. Center of the Universe Come June, the Bay Area becomes the beating heart of the data and AI world, and Chalk will be right in the middle of it. Snowflake Summit (June 1–4) This is our first year as an official sponsor, and we're not phoning it in. We've also been heads-down for months on a set of product announcements we're saving for Summit. Big things. That's all I'll say for now. The week of, you'll see. Snowflake Summit is one of the biggest GTM moments of the year for us. A lot of our customers already run Snowflake as their offline store and Chalk as the real-time inference layer on top. Snowflake is the analytical and training backbone. Chalk is what powers production AI on top of it: features computed in milliseconds, federated queries against live data, point-in-time correct training sets, sub-5ms inference. If you've heard us talk about that architecture, this is where we plant the flag in person. Here's where to find us all week. Booth #2611, Moscone South. Live demos with our engineering team every day of the conference. Real code, real queries, real latency. If you've been curious about how teams compute features in milliseconds against Snowflake data, come watch us do it. Two speaking sessions from our co-founders. The first is a 45-minute customer talk: "The Infrastructure Behind Grindr's AI Transformation" with Marc Freed-Finnegan (CEO, Chalk) and AJ Balance (CPO, Grindr). It's the story of how Grindr rebuilt their AI infrastructure, what changed, what they learned, and what's now possible that wasn't before. The kind of customer talk we don't usually get to give in public. The second is a 20-minute technical deep dive from Elliot Marx, our co-founder, on building a real-time feature store in the Snowflake ecosystem. Federated queries, point-in-time correctness, sub-5ms inference, and why the gap between offline and online is still the hardest problem in production ML. Kickoff Party, evening of June 1. Drinks, food, our team, our customers, the people building real-time AI on Snowflake, all in one room. Invite-only. Send me a note if you want on the list. Time with our co-founders and exec team. Marc, Elliot, Andy, and the rest of our leadership will all be at Summit and available for 1:1s with technical and business leaders throughout the week. If you're evaluating a feature store, replatforming off legacy infrastructure, or trying to figure out whether Chalk fits your stack, this is the most concentrated chance you'll have to talk to us in one place. If you're building AI on Snowflake, this is the conference to find us. We'll be there all four days. Databricks Data & AI Summit (June 15–18) Two weeks later. We have something cooking. For teams that need to go fast. That's all I'll say. You'll enjoy it. The Big Picture Spring Fling isn't just a marketing calendar. It's a statement. Chalk just had its best pipeline quarter ever, our team has doubled in the last year, and the category around real-time AI infrastructure is moving fast. We're not slowing down. We're planting a flag. If you're heading to any of these events, let's find time. I'll be at all of them. Shoot me a note at alexandra@chalk.ai, or grab a coffee, drink, dinner, whatever's easy. See you out there. # What Is Data Drift in Machine Learning | The Chalkboard source: https://chalk.ai/blog/data-drift Learn what data drift is in machine learning, how it differs from concept and model drift, and how teams detect and manage drift in production ML systems. Intro Your model passed every test before deployment. A month later, fraud is slipping through. Recommendations are missing. Credit approvals are off. Nothing in your codebase changed. The model is still running - but something is wrong. The culprit is often data drift. Data drift occurs when the real world data diverges from the data a model was trained on. It’s not a bug or an operational mistake—it’s an inherent property of production machine learning. The inputs your model sees today are not the same as they were six months ago, and most models have no built-in way to adapt.The result is silent degradation: a widening gap between what the model expects and what it actually receives. Over time, that gap erodes performance until predictions no longer reflect reality. This guide covers: - What data drift is (and isn’t) - How it relates to covariate shift, concept drift, and other forms of distribution shift - Common causes of drift in production systems - How to detect it early—before it impacts users - How modern ML teams manage drift at scale What Is Data Drift? Data drift (also known as covariate shift) occurs when the statistical distribution of input features — P(X) — in production changes relative to the distribution a model was trained on. The model itself hasn't changed. Its weights, logic, and learned patterns are all identical - but the data flowing through it at inference time no longer resembles the data it was optimized for. Therefore, predictions become progressively less reliable. Data drift is not noise. Random fluctuation is normal and expected. Data drift is systematic: a sustained, directional shift that persists over time and reflects a real change in the environment a model operates in. Some examples include: - Fraud detection: A model trained on 2023 transaction data begins seeing unfamiliar patterns—new device types, payment methods, and geographies—as the user base evolves. - Credit risk: A model trained primarily on salaried employees is applied to a growing percentage of gig workers, whose income and spending patterns differ structurally. Recommendations: A model trained on pre-pandemic behavior sees shifts in available inventory and user interaction patterns, changing the distribution of inputs it receives. In each case, the underlying data has changed in a way that invalidates the model's assumptions. The model has no built-in awareness of this shift - it keeps making predictions. They're just increasingly wrong. Data Drift and Related Concepts Data drift is often used as a catch-all for any kind of model degradation in production. But several distinct failure modes look similar from the outside while requiring very different interventions. Recall that data drift is specifically a change in P(X) — the distribution of model inputs. Each concept below breaks something different, and knowing which one you're dealing with determines how you respond. Data Drift vs. Concept Drift Data drift is a change in inputs. Concept drift is a change in rules. Formally, concept drift describes a change in P(Y|X) — the relationship between inputs and the outcome you're trying to predict. The inputs may look the same, but what they mean has changed. A concrete example: in a fraud detection system, data drift shows up as new transaction patterns — more mobile payments, more cross-border activity — that the model wasn't trained on. Concept drift looks different: the transaction mix is familiar, but fraudsters are using a new attack vector, so patterns that once signaled legitimate activity now correlate with fraud. The model's decision boundary is no longer valid. Concept drift typically can't be solved by retraining alone — it often requires rethinking the features themselves and sometimes the model architecture. Data Drift v. Label Shift Data drift is a change in what the model sees. Label shift is a change in what actually happens. Formally, label shift (also called prior probability shift) describes a change in P(Y) — the distribution of outcomes. The class balance shifts even if the relationship between features and labels remains similar. A concrete example: a fraud model is trained and calibrated against a 0.3% fraud rate. An economic downturn drives more people toward fraudulent behavior — the same types of fraud, through the same channels, just more of it. The feature distributions haven't changed and the fraud patterns look familiar. But the base rate the model was tuned to no longer holds — fraud is now at 1%, so its scores are miscalibrated, thresholds are wrong, and alert volumes are off. Label shift means the model's calibration is broken — the fix is recalibrating thresholds, adjusting decision rules, and reviewing any downstream logic built on top of model scores. It often demands a faster response than data drift because it directly affects operational parameters — alert volumes, approval rates, and risk thresholds can all break before anyone notices the underlying distribution has moved. Data Drift vs. Feature Drift Data drift is the aggregate picture. Feature drift is the per-feature signal. Formally, feature drift describes changes in individual feature distributions over time, sometimes in ways that aren't visible from aggregate statistics. A single feature can drift significantly while the broader distribution looks stable, and a model can degrade from one corrupted input while dozens of others remain clean. A concrete example: a fraud model uses 40 features. Aggregate monitoring shows stable overall distributions. But one feature — a 30-day rolling average of transaction count — has quietly shifted because of a seasonal pattern the training data didn't capture. The model's accuracy starts slipping, and aggregate drift metrics don't flag anything because the shift is isolated to one input. Feature-level monitoring catches it immediately. This distinction matters because feature drift is the earliest detectable signal of broader distribution problems. In most cases, individual feature distributions shift before model accuracy degrades. By the time a drop in accuracy is measurable, the feature drift that caused it may have been present for weeks. Teams that monitor at the feature level — tracking distributions, freshness, and statistical drift per feature, not just per model — get earlier warnings and more actionable diagnostics: they know exactly which feature shifted, not just that something went wrong. Data Drift vs. Model Drift Data drift is a cause. Model drift is the symptom. Formally, model drift describes the downstream result: the model's performance is degrading in production — accuracy drops, predictions worsen, business metrics decline. It isn't a distinct type of distribution shift; it's the observable effect of one. Data drift is one of the most common root causes, but not the only one. A concrete example: a credit model's approval-to-default ratio starts climbing. That's model drift — something is wrong with the model's decisions. But the cause might be data drift (the applicant population shifted), concept drift (the relationship between borrower features and default risk changed), or even a pipeline issue (a feature started arriving stale). The degradation looks the same from the outside; the fix depends entirely on which upstream problem caused it. This distinction matters for how you respond. If you observe model drift and jump straight to retraining, you're treating the symptom — not fixing the underlying issue. If your feature pipeline has changed, or user behavior has shifted in ways that invalidate your training distribution, retraining on the same assumptions just reintroduces the problem, often faster. Effective model drift response starts with diagnosing which form of drift caused it, not just that performance dropped. Data Drift vs. Training–Serving Skew Data drift is a change over time. Training–serving skew is a mismatch across environments. Formally, training–serving skew describes an environmental mismatch: training and inference compute different values for the same features — even at the same point in time. Where data drift means the world changed, training–serving skew means your pipelines diverged. The model isn't seeing a shifted world; it's seeing a differently computed version of the same world. A concrete example makes the distinction clearer. A model is trained using a batch-computed feature — say, a 30-day average transaction value calculated in a data warehouse overnight. At serving time, the "same" feature is computed from a real-time API that only has access to the past 7 days of transactions. The feature is nominally the same, but represents different windows of data. The model is now operating on inputs it was never trained to handle — not because the world changed, but because the feature pipelines diverged. The two problems compound each other. If your training pipeline handles timestamps differently than your serving pipeline, even a small distribution shift in the underlying data can cause unexpectedly large degradation as the features themselves are inconsistently computed. Systems that enforce point-in-time correctness — ensuring that training and serving use identical feature logic evaluated at the same moment in time — eliminate training–serving skew entirely, and make data drift easier to isolate when it does occur. dataset = ChalkClient().offline_query( input={ # Sample business 123 twice because we have two input_times Business.id: [123, 123], }, # Each element of `input_times` corresponds to the element # with the same index in `input` input_times=[t1, t2], # Sample all features of business. # Alternatively, sample only the features you need: # output=[Business.sales, Business.cogs] output=[Business], ) python Why Data Drift Matters in Production ML The most dangerous thing about data drift is how quiet it is. There’s no error in the logs. No pipeline failure. No alert fires. The model keeps running, keeps producing outputs, and often keeps looking healthy by the metrics teams typically watch. Meanwhile, the decision quality is eroding. By the time data drift becomes visible through model accuracy metrics or business KPIs, it’s usually been present for weeks. Feedback loops are slow; you may not know whether a fraud decision was correct until a chargeback arrives 30 days later, or whether a credit approval was sound until payment behavior develops over months. In that window, a drifted model may make thousands of degraded decisions, and you only find out when the downstream numbers start to move. The consequences depend on the use case, but they’re never trivial: - In fraud detection, a drifted model lets more fraudulent transactions through, or flags more legitimate ones, eroding both revenue and user experience. - In credit decisioning, degraded risk scores lead to mispriced loans, increased default rates, and regulatory exposure. - In recommendation systems, stale feature distributions produce irrelevant recommendations, reducing engagement, conversion, and trust. There are also audit and compliance dimensions. In regulated industries like fintech, healthcare, and insurance, teams need to be able to reproduce historical model predictions and explain why a specific decision was made. If your feature pipeline didn’t capture what the model actually saw at inference time, that reproducibility is gone. And these challenges aren't limited to traditional ML — LLM-based systems face analogous risks when the prompt inputs or retrieval context they depend on shift over time. Common Causes of Data Drift Drift rarely appears out of nowhere. It has causes, and those causes fall into recognizable patterns. Identifying which category you’re dealing with determines both how you detect it and how you respond. Changes in User or System Behavior The most common cause of data drift is simply that the world changed. Users behave differently over time. New product lines change the composition of the customer base. Regulatory changes alter how people interact with financial products. The tricky part is that successful product decisions can cause drift. When a fintech app adds a new payment product — say, instant transfers or buy-now-pay-later — the transaction distribution changes. Users who adopt the new product have different behavioral signatures than the original base. A fraud model trained before the launch may not have learned the right patterns for the new product, and a credit model trained on the old user mix may not price risk correctly for the expanded population. Seasonality and Cyclical Patterns Seasonal drift is predictable in hindsight but catches models off guard in practice. Most models are trained on a fixed window of historical data that may not span a full seasonal cycle — or may overrepresent one season relative to the conditions the model faces in production. A concrete example: a credit model trained primarily on Q1–Q3 data sees a significant shift in spending patterns during the holiday season. Transaction volumes spike, average purchase values increase, and the mix of merchant categories changes. None of this represents a permanent shift in the population — it's cyclical. But from the model's perspective, the input distribution has diverged from training, and predictions degrade until the cycle passes or the model adapts. Seasonality is particularly insidious for time-windowed features. A 30-day rolling average of transaction count will look dramatically different in December versus February, even for the same user. If the training data didn't capture this range, the model treats seasonal behavior as anomalous. Teams in fintech, e-commerce, and insurance — where seasonal patterns are pronounced — often need separate monitoring thresholds for known cyclical periods, or training datasets that deliberately span full annual cycles. External Events Macro events can shift data distributions faster than any scheduled retraining cycle can adapt. A fraud wave using a new synthetic identity technique invalidates KYC feature baselines overnight. An economic shock changes the income distribution of loan applicants. A regulatory change shifts what data is available or how it’s reported. External events are particularly challenging because they’re often sudden and unpredictable. The monitoring strategies that catch gradual drift — tracking distributions over rolling windows — may not detect a rapid shift until several days of data have accumulated, by which point the model has already been making poor decisions throughout the event. How to Detect Data Drift Effective drift detection is not a periodic audit — it’s continuous monitoring designed to surface signals before they become failures. The most important architectural choice is where you instrument: teams that monitor at the model output layer see degradation after it’s already happened. Teams that monitor at the feature layer see it coming. Summary Statistics and Visual Monitoring The most accessible form of drift detection is tracking basic summary statistics per feature over time: mean, median, standard deviation, null rate, min/max, and value distributions. When any of these deviate significantly from a reference window — typically the training dataset or a recent stable production window — it’s a signal worth investigating. Visual monitoring — histograms, distribution overlays, and time-series charts of feature means — is useful for debugging and root cause analysis but difficult to scale. With hundreds or thousands of features across multiple models, manual inspection can’t keep up with the pace of production traffic. Summary statistics are a good foundation; they need automated alerting to be operationally useful. Statistical Tests Statistical tests provide a more rigorous signal than raw summary statistics. The Kolmogorov-Smirnov (KS) test measures the maximum difference between two cumulative distribution functions — typically the training distribution and the current production distribution — to detect whether they’re drawn from the same underlying distribution. When the test statistic exceeds a critical value, it indicates statistically significant drift. The Population Stability Index (PSI) takes a different approach: it quantifies how much a distribution has shifted relative to a reference, producing a single number that scales with drift severity. PSI scores below ~0.1 typically indicate stable distributions; scores above ~0.25 commonly signal significant drift requiring action. These are widely used rules of thumb, though appropriate thresholds vary depending on the model, use case, and how fast the data environment changes.Other commonly used methods include the Chi-Squared test for detecting shifts in categorical feature distributions, Jensen-Shannon divergence for a bounded and symmetric measure of distribution distance, and Wasserstein distance (also known as Earth Mover's Distance) for capturing the magnitude of distribution shifts in a way that accounts for the geometry of the feature space. All these tests involve tradeoffs. More sensitive configurations detect drift earlier but generate more false positives. Less sensitive configurations reduce noise but risk missing early-stage drift. The right calibration depends on how fast your data environment changes, how costly false alarms are, and how much lead time you need before drift affects model quality. Feature-Level Monitoring The statistical tests above are most powerful when applied at the feature layer rather than the model output layer. As covered in the feature drift section earlier, aggregate monitoring can miss isolated feature shifts entirely — and by the time accuracy drops are measurable, the underlying drift has often been present for weeks. In practice, feature-level monitoring means treating each feature as an independently observable signal. You track its distribution against a reference window, its freshness — the gap between when it was computed and when it's used at inference time — and its drift test statistics over time. When a feature crosses a threshold, you know exactly which part of your data pipeline to investigate before you ever see a model quality alert. That specificity is what cuts diagnosis time from days to hours. Modern feature infrastructure makes this kind of granular monitoring practical at scale. Rather than building and maintaining separate observability tooling, teams can instrument at the feature definition level and get drift alerts, access pattern tracking, and freshness monitoring as built-in properties of how features are computed and served. How to Handle and Mitigate Data Drift Once drift is detected, the harder question is what to do about it. The right intervention depends on which features are affected and what the business impact is. Root Cause Analysis with Lineage The first step is diagnosis, and lineage is the tool that makes it possible. Data lineage — a traceable record of where each feature comes from, how it’s transformed, and which models depend on it — lets teams trace a drift signal back to its source. Without lineage, a model accuracy drop sends teams into a debugging process that can take days: checking pipelines, reviewing upstream sources, testing hypotheses about which feature might have changed and why. With lineage, a drift alert on a specific feature includes its provenance — which upstream table it came from, which transformation produced it, which other models use it. Debugging time can drop from days to hours. Lineage also defines the scope of impact: when a feature drifts, how many models are affected? Which downstream decisions should be reviewed? This is the kind of question that’s nearly impossible to answer reliably without a system that tracks feature dependencies end to end. Retraining Strategies Retraining is the most common response to data drift, and it’s often the right one — but it’s not a catch-all. There are two broad approaches: scheduled retraining, where models are retrained on a fixed cadence regardless of detected drift; and trigger-based retraining, where retraining is initiated when drift signals exceed a threshold. Trigger-based retraining is generally more efficient but requires solid monitoring infrastructure to be reliable. Scheduled retraining is simpler to operate but may retrain models unnecessarily when there’s no meaningful drift, or fail to retrain quickly enough when drift is sudden. The most important caution: don’t retrain without understanding why the model drifted. If a data pipeline is producing corrupted features, retraining on that pipeline’s output trains a model on corrupted data. The retrained model inherits the problem. Before triggering a retrain, confirm that the training data for the new run reflects the intended distribution — not the drifted one. Feature Engineering Adjustments Some drift doesn’t call for retraining at all; it calls for adjusting how features are defined. Time windows are the most common lever: if user behavior has accelerated, a 30-day rolling average may be too slow to capture meaningful signal, and shortening to a 7-day window may restore model performance without a retrain. Other adjustments include normalization strategy changes (if the scale of raw values has shifted), decay weighting (to down-weight older data that no longer represents current behavior), and feature retirement (when a feature has become so unstable that its variance introduces more noise than signal). Process and Governance Changes Drift is a technical problem with an organizational dimension. Technical tooling can detect it and surface it, but responding effectively requires clear ownership: who is responsible for monitoring drift on a given feature set? Who gets alerted when a threshold is crossed? Who has authority to initiate a retrain or a feature change? In regulated industries, the governance layer isn’t optional. The ability to demonstrate that drift was detected, investigated, and resolved within a defined timeframe — with documentation of what the model saw and when — is increasingly a compliance requirement in credit, insurance, and healthcare ML systems. Data Drift in Real-Time ML Systems Real-time ML systems experience drift differently than batch systems. Feedback loops are shorter. Behavioral shifts propagate faster. And unlike batch inference, where a delay of hours may be acceptable, real-time decisions require features that reflect what’s true right now — not what was true when the last batch job ran. This creates a specific set of challenges: - Time-aware features — rolling windows, recency-weighted aggregations, behavioral velocity metrics — are among the most predictive features in real-time systems and among the most vulnerable to drift. When user behavior accelerates or seasonal patterns shift, these features move fast. - Freshness guarantees matter at a different level of precision. A feature that’s two hours stale may be acceptable for a weekly batch scoring job. For a fraud decision made in under 100ms at the point of transaction, a two-hour-old velocity feature can be the difference between catching fraud and missing it. - Reproducibility is operationally critical. When a real-time model makes a decision that needs to be audited or challenged, teams need to reconstruct exactly what the model saw at that moment — which features, which values, which version. Systems that don’t capture inference-time feature snapshots make this impossible. Often, the best answer is to compute features at inference time from live sources rather than serving pre-materialized values from a store that’s only as fresh as the last batch run. This eliminates the staleness ceiling that batch materialization creates and ensures that the features a model uses at decision time reflect current reality. Data drift FAQs A few questions teams have once a model is live and accuracy starts slipping. What is data drift? Data drift is when the live data a model sees in production moves away from the data it was trained on. It is a normal property of production ML, not a bug. It shows up as a slow decline in accuracy as the gap between expected and actual inputs widens. What is the difference between data drift and training-serving skew? Data drift is a mismatch over time: the world changes, so today's inputs differ from the training data. Training-serving skew is a mismatch at a single moment, where training and production compute different values for the same feature even at the same point in time. How do you catch data drift before it hurts a model? Monitor at the feature layer rather than only the model's output, keep lineage from each feature back to its source, and serve features that reflect current data instead of stale batch values. Being able to reproduce what a model saw at any past moment makes drift far easier to trace and fix. Drift Is Inevitable. Chalk Can Help. Data drift is not a problem you solve once and move on from. It’s a permanent feature of deploying machine learning in a world that keeps changing. The question isn’t how to prevent drift — you can’t. The question is how to build systems that detect it early, diagnose it fast, and respond before it causes failures. The teams that handle drift best share a few structural properties. They monitor at the feature layer, not just the model output layer, so they see signals before accuracy degrades. They maintain data lineage end to end, so when a feature drifts they can trace it to the source and scope the impact. They enforce freshness — their models get features that reflect current reality, not a stale cached value from a batch pipeline that ran twelve hours ago. And in real-time systems, they can reproduce exactly what a model saw at any point in time — not just for audits, but for debugging and retraining. Chalk is built around this architecture. Mission Lane uses Chalk to power real-time credit approvals, fraud detection, and customer-facing features from a single consistent feature platform — and credits it with helping the team reduce drift and scale confidently across systems. Verisoul ships fraud detection updates 10x faster and achieves 4x more accurate detection using fresh inference-time features, with complete auditability for every decision. Across both cases, the architectural foundation is the same: feature-level observability, lineage, and freshness-aware compute, treated as system-level properties rather than afterthoughts. “It used to take 24 hours to regenerate a training set. Now it’s under an hour. We can trust every feature in it.” — Robert Theed, Backend Tech Lead, iwoca Modern MLOps maturity means designing systems that expect drift, not systems that assume stability. That means feature-level observability that catches distribution shifts before they cascade. It means lineage that makes root cause analysis fast. It means freshness guarantees that ensure your models are making decisions on current reality, not a snapshot from last week. And increasingly, it means extending those same properties into LLM pipelines, where prompt inputs can drift just as quietly as traditional feature distributions. Drift is inevitable. Failures don’t have to be. See how teams move from reactive drift detection to feature-level observability with Chalk. [Talk to an engineer] → Get a Demo # Why Your Feature Store Has a Freshness Ceiling | The Chalkboard source: https://chalk.ai/blog/why-feature-stores-have-freshness-ceiling Most feature stores are built around batch pipelines and materialized storage, which creates an inherent freshness ceiling. Tightening staleness settings cannot overcome architectural limits. This article explains why traditional feature stores struggle with real-time freshness and how a compute-first approach enables on-demand feature computation, lower latency, and improved correctness for time-sensitive ML use cases like fraud detection and underwriting. Intro A fraud model has milliseconds to decide whether something is legitimate. A signup could be genuine or automated. A transaction could be a real purchase or a stolen card. To do so, the model pulls together signals: device fingerprints, behavioral patterns, network risk data. In Machine Learning systems, these signals are features. Features are typically managed by a feature store: the infrastructure that computes, stores, and serves them to models in production. But the signals that features represent change constantly. A device fingerprint that was valid an hour ago may no longer reflect the current session. What is feature freshness? Feature freshness is that gap between when the underlying data changes and when an updated feature is available at inference time. In contexts like fraud detection, even 100 milliseconds of staleness can mean the difference between catching a fake account and letting it through. Feature freshness directly impacts the quality of predictions in systems that rely on real-time signals. Examples of freshness-sensitive ML applications include: - Fraud detection - Credit underwriting - Abuse detection - Recommendation systems - Ad targeting Freshness as an architecture problem Most teams, when they notice stale features, look at the serving layer first. They tighten freshness windows, shorten sync intervals between offline and online stores, and run batch pipelines more frequently. These are reasonable instincts since the serving layer is the most visible and most tunable part of the system. But this treats freshness as a configuration problem: a matter of tuning settings without examining the assumptions baked into the architecture itself. If tuning settings do not change the value your model sees, your architecture has a freshness ceiling. To understand why those fixes hit a ceiling, let's look at how most feature stores actually work. Where traditional feature stores fall short on freshness In a typical ML system, getting a feature to a model involves a long chain of steps. Raw data gets extracted from source systems, transformed, aggregated, and written to an offline store. A batch pipeline, usually running on a schedule (say hourly or daily), handles this processing. The results are then synced to an online store, which is what the model actually reads from at inference time. The feature store, in most implementations, sits at the end of this chain. It's primarily a storage and serving layer. It serves pre-computed values from upstream systems. This is a storage-first, or cache-first, architecture. It carries a set of assumptions: - All data must be moved to one place before features can be computed. - Features are computed in batch jobs, not at serving time. - Values must be materialized and stored before they can be served. - Training and inference run on separate infrastructure, with feature logic often rewritten between the two. These assumptions create limitations beyond the scope of tuning. 1. There is a ceiling on freshness you cannot exceed. The online store can only serve what has already been materialized from the most recent pipeline run. Tightening staleness settings does not make the value fresher. It simply re-reads the same data. If the upstream batch has not run, the feature cannot reflect new changes. The architecture defines the freshness ceiling. 1. Pushing freshness lower means escalating complexity. When serving-layer fixes aren't enough, teams move upstream: shorter batch intervals, streaming pipelines layered onto batch, synchronization logic, and monitoring drift between online and offline stores. Each workaround adds cost and operational burden. What freshness requires To minimize the lag between when source data changes and when updated features are available at inference, you need to move beyond tuning the last mile: - Compute where the data lives. Rather than waiting for data to be extracted, staged, and moved through a pipeline before any computation can begin, compute features at the source. - Support on-demand computation. The system should be able to execute feature logic at query time, rather than only returning whatever was last written to the store. - Treat materialization as an optimization, not a requirement. Cache when it helps latency. Do not depend on it for correctness. This is the shift from storage-first to compute-first architecture. In a compute-first system, freshness is bounded by source latency and execution time, not by materialization schedules. Tradeoffs When designing a feature store for real-time ML, there are a few tradeoffs worth thinking through. Materialized vs. on-demand Pre-computing features can be faster to serve but inherently stale. On-demand computation is fresh but adds compute at query time. Most systems force you to choose this tradeoff at the architecture level. A compute-first system allows you to choose per feature. A transaction sum can be materialized into time-bucketed aggregations for performance, while a continuous buffer fills in the gap between the last backfill and now with live data. Latency vs. correctness Serving a cached value is fast. Serving the right value may take a few more milliseconds. In domains like fraud, abuse, and underwriting, correctness wins. Batch, streaming, and real-time Some features don't need real-time freshness. A 30-day aggregate can be batch-computed daily. Others need to reflect changes within minutes, making them natural fits for streaming. And some must be computed at the moment of inference. The architecture should support all three without requiring separate systems or duplicated logic. How Chalk approaches freshness Chalk is a feature store built around computation, sometimes also referred to as a feature engine. Where traditional feature stores are storage-first, Chalk is compute-first. It stores, serves, and reuses features across training and production. But it can also compute features on demand at query time. When you query a feature, Chalk can execute the function directly, traverse dependencies, fetch fresh data, and run transformations on the fly. Several architectural choices make this possible. Source-agnostic. Chalk connects directly to underlying data sources, whether that's a database, an API, or a Kafka stream. Data doesn't need to be staged or pre-loaded before features can be computed. Federate. Chalk fetches data where it lives instead of requiring centralization. A lending decision might combine internal transaction history with real-time income data from Plaid and bureau data from TransUnion in a single query. Unified online and offline definitions. The same Python feature definitions run in both training and production contexts. This eliminates training-serving skew by design. On-demand computation. Chalk's query planner dynamically builds an optimized execution plan based on feature dependencies and available sources at query time, rather than relying on what was last written to the store. Materialization when it helps, on-demand when it matters. Chalk can pre-compute and cache where latency requires it, while always retaining the ability to compute fresh values when correctness demands it. The result is a system where materialization is an optimization, not the source of truth. transaction_count: Windowed[float] = windowed( "30m", "6d", "12h", materialization={"bucket_duration":"10m"}, expression=_.transactions[ _.timestamp <= _.chalk_now, _.timestamp > _.chalk_window ].count(), ) python For example, a rolling 30-minute transaction count can be partially materialized for performance, with streaming data from Kafka updating the aggregation as new events arrive. The full path from source data to served feature runs in single-digit milliseconds, even across heterogeneous sources. Where computation happens defines your freshness ceiling When computation and serving are unified rather than separated by layers of ETL, freshness stops being something you fight for and becomes something the architecture enables by default. If tightening your freshness window or lowering staleness does not change your feature values, your architecture is telling you something. If you are evaluating how your feature platform handles real-time decisions, start by asking: Where does computation happen? If the answer is “upstream in a batch job,” you already know where your freshness ceiling lives. To go deeper into how compute-first feature systems work in practice, explore our Architectures and reach out to our FDEs. # Quarterly Product Update: Winter source: https://chalk.ai/blog/product-update-feb-2026 A roundup of Winter 2025’s latest Chalk product updates, covering performance improvements, new integrations, and infrastructure enhancements for running fast, reliable real time data pipelines in production. This quarter, we focused on strengthening the systems that keep Chalk fast and reliable in production. Updates to planning, monitoring, and integrations improve how teams scale heavy workloads, automate alerts, and manage real-time data pipelines. Here’s what’s new. Metaplanning and autosharding for large workloads Chalk’s metaplanner now automatically determines how to shard scheduled offline queries based on input size and complexity. It splits large inputs into multiple smaller queries that the query planner can execute in parallel across available compute. Previously, teams needed to decide shard counts ahead of time. The metaplanner now computes the optimal number of shards automatically, scaling compute to match workload size. This makes feature recomputations faster, more efficient, and easier to manage. It improves runtime behavior and reduces operational overhead across high-volume jobs. Get in contact with our team to enable metaplanning in your environment. Runtime and planner performance improvements We improved planner caching and parallel resolver execution to make query runtime faster and more consistent. Cold-start latency has been reduced, and repeated workloads now benefit from higher cache hit rates. These changes help teams achieve lower latency and smoother performance during both development and production runs, especially when operating at scale. Webhooks for real-time alerting and custom integrations Teams can now configure webhooks directly in the Chalk dashboard to send alerts or trigger actions when events occur. Webhooks can be set up for scheduled queries, task completions, or deployment failures, and can integrate with tools like Slack, PagerDuty, and custom HTTP endpoints. These webhook updates are part of broader monitoring improvements that include configurable alert rules and metric filters, helping teams fine-tune their observability and catch issues earlier. This helps teams automate responses and maintain system reliability across environments. Learn more about Chalk’s latest webhooks. Connect vector databases from the dashboard Chalk now supports registering external vector databases directly through the Integrations page. Once connected, teams can use these stores to power retrieval, personalization, and fraud detection models without leaving Chalk. The dashboard supports registration, configuration, and health monitoring of the integration. This reduces setup complexity and expands Chalk’s ecosystem for embedding and similarity-based systems. Learning and deep-dive resources To help teams see how Chalk fits into production ML stacks, we partnered with customers to share their stories and expanded our content across the website and docs. - New use case pages: Explore how Chalk powers real-time decisioning across different use cases. - Fraud detection - Detect and prevent fake accounts and transactions using live features. - Payments - Power authorization and risk decisions with fresh behavioral data. - Underwriting - Deliver auditable credit decisions with fresh bureau and alternative data. - Recommender Systems - Serve real-time recommendations that adapt instantly to user behavior. - Search and Ranking - Rank results dynamically at query time using living context. - Growth Decisioning - Power lifecycle and retention decisions with live engagement data. - Chalk by team pages: Role-specific content and highlights for teams building and operating production ML: - Data Engineers - Build real-time features without maintaining custom pipelines. - MLOps - Operate production ML with full data lineage and reliable model inputs. - ML Engineers - Ship low-latency inference systems that serve fresh data at scale. - Data Scientists - Prototype and deploy features from notebooks with parity across training and serving. - Customer story: We partnered with Medely to showcase how they use Chalk to power clinician matching in real time, optimizing for availability, credentials, and location across millions of jobs and professionals. - Migration guide: We published a migration guide for teams moving from Tecton to Chalk. It walks through how to bring your existing features and pipelines into Chalk’s feature store platform. Learn how Chalk supports transitions from Tecton here. # Chalk x ODSC Meetup 2026 source: https://chalk.ai/blog/chalk-odsc-meetup-recap-2026 Recap from ODSC's first SF meetup of 2026, where we did a deep dive into real-time feature computation for marketplace recommender systems. Last week, we kicked off 2026 with ODSC's first San Francisco meetup of the year at our new Union Square office! We had ML engineers from fintech, e-commerce, and enterprise software come through to hear our co-founder Elliot Marx talk about why most marketplace recommender systems fail in production, and what to do about it. Spoiler: it's not a model quality problem. The Core Problem Most production recommenders rely on batch pipelines and precomputed features. For marketplaces, this creates two ceilings on relevance: you can't precompute every user-item combination, so you have to guess which ones to score. And even for the pairs you do precompute, the features go stale before decision time. You're either missing users or serving them outdated context, often both. The Live Demo Elliot's talk centered on a different approach: computing features on demand in real-time. Think of it less like a cache-first feature store and more like a query execution engine: determining at request time which resolvers to run, which data sources to hit, and how to join the results under tight latency budgets. Elliot's demo walked through the mechanics: Writing resolvers: He showed how to define new features in Python, like a `name_email_match_score` that computes similarity between a user's name and email. Resolvers declare their dependencies in function signatures, making the feature graph explicit. Query planning and execution: When a feature is requested, Chalk generates a query plan and determines the optimal execution path. It automatically pushes filters down to your data sources and compiles Python resolver logic to C++ to hit single-digit millisecond latencies at production scale. Joining across data sources: The system queries across Snowflake, Postgres, and online stores at request time, so you don't need to denormalize everything upfront. Temporal consistency: For training data, Chalk performs point-in-time lookups so your model never trains on features that wouldn't have existed at prediction time, preventing data leakage when backfilling new features. The Q&A ran well past schedule: engineers stuck around with drinks and bites, digging into query optimization, handling highly dynamic features, and how this architecture compares to their current setups. What Resonated The room was full of people working on similar problems: keeping features fresh under tight latency constraints and dealing with multi-entity relationships in marketplaces. When you treat recommender systems as query execution problems instead of offline prediction pipelines, this means: - No more guessing which user-item pairs to precompute - No more stale features because batch jobs run on a schedule - One codebase for training and serving - Richer cross-entity signals because you can join at decision time instead of flattening relationships upfront For teams running marketplace recommenders, fraud systems, or any high-frequency decision system where context changes constantly, this reframing changes what's architecturally possible! What's Next This was the first of many events we're planning in our SF office. If you're working on production ML systems and want to see Chalk in action, book a demo to walk through how decision-time feature computation could work for your use case. Thanks to everyone who came through, and to ODSC for co-hosting a great start to the year! # 7 Common MLOps Challenges source: https://chalk.ai/blog/7-common-mlops-challenges MLOps enables teams to train and run models at scale. Learn why production ML is challenging, from feature inconsistency to fragmented data and how to solve it. MLOps, short for Machine Learning Operations, refers to the work of enabling teams to train and run models at scale. It covers the systems and practices that transform experiments into models running in production: how models are deployed, how data and features stay consistent over time, and how performance is monitored. While these challenges show up differently depending on role, they tend to surface across the entire ML lifecycle: from experimentation to deployment to long-term maintenance. Data scientists feel them when experiments don’t translate cleanly to production, data and ML engineers feel them when infrastructure becomes brittle or slow, and MLOps teams feel them when reliability, governance, and observability break down at scale. Why Is MLOps So Challenging? Such work is difficult not because teams lack expertise, but because machine learning behaves fundamentally differently from traditional software. In standard applications, behavior changes only when engineers modify the code. In ML systems, behavior can change even when the code stays the same, because the data feeding the model is constantly evolving. User behavior shifts, upstream systems change, new edge cases appear, and distributions drift. As a result, production ML is not a one-time deployment, but a continuous process of adaptation. While software engineering has had decades to converge on shared deployment patterns and tooling, MLOps is still a relatively young discipline. Many organizations are building their ML systems by stitching together tools that were never designed to work as a cohesive whole. The result is a recurring set of challenges that slow teams down and undermine model reliability. Common MLOps Challenges 1. Fragmented Data + Loss of Traceability Most production models rely on data from many different systems at once. Historical information typically lives in a data warehouse such as Snowflake or BigQuery, while real-time state comes from operational databases like Postgres or MySQL. Fast lookups may be served from key-value stores such as Redis or DynamoDB, and additional context is often pulled from third-party APIs for predictions like risk scoring, identity verification, or pricing. Each of these systems updates on a different schedule and is owned by a different team. When a model’s predictions start to drift or degrade, it can be extremely difficult to determine which input changed or where the issue originated. Without clear data lineage and traceability, teams are unable to identify the root causes of model drift right away, often spending days debugging. This lack of end-to-end traceability is why many teams look for a single system of record for features: one that can track where features come from, how they’re computed, and which models depend on them across both historical and real-time data. Without that shared foundation, observability and auditability are perpetually bolted on after the fact. 2. Feature Inconsistency A feature can be an input or an output to an ML model, often derived by transforming raw data into a more useful signal through feature engineering. In many organizations, features are computed one way during model training and another way when the model runs in production. (If you're newer to this concept, see What is a Feature Store for a full breakdown.) For example, a user attribute might be backfilled in a warehouse for training purposes but be missing or delayed in a real-time operational database at serving time. Batch pipelines may handle timestamps or missing values differently than real-time systems. In other cases, engineers rewrite feature logic entirely when moving from notebooks to production APIs. These differences rarely cause explicit failures. Instead, they introduce subtle inconsistencies that quietly degrade model performance over time: a phenomenon also known as train-serve skew. This is common when responsibility blurs between data scientists defining features in notebooks and engineers rebuilding them for production systems, each with good intentions but no shared execution layer. Eliminating this class of problems requires defining features once and reusing them everywhere, rather than relying on parallel pipelines and conventions to keep systems in sync. 3. Disconnected Training + Serving Environments Training and serving ML models usually happen in different environments, with different assumptions about data. During training, models typically learn from historical data stored in data warehouses. Once deployed, those same models are expected to run inside production systems, executing on live, real-time data pulled from operational databases, APIs, caches, or streaming platforms like Kafka, often under much stricter latency and reliability constraints. For ML engineers, this gap between experimentation and serving is where velocity is lost. Models ship, but confidence erodes once they hit production traffic. Each layer of this stack has its own dependencies, configurations, and failure modes. Small mismatches—such as differences in library versions or transformation logic—can lead to incorrect predictions without triggering errors. Because the system technically “works,” these issues can persist unnoticed for long periods. Eliminating this class of problems requires a feature store where features are defined once and reused everywhere. 4. Monitoring Systems Without Understanding Model Behavior Most production monitoring focuses on infrastructure metrics like uptime and latency. While these signals are necessary, they are not sufficient for machine learning systems. A model can be serving predictions quickly and reliably while still producing worse outcomes. This often happens when a model’s inputs change in subtle ways. A third-party API may alter a field format, a streaming pipeline might drop events, or a feature sourced from a cache could start returning null values. Without visibility into feature values, distributions, freshness, and provenance, teams often end up scrutinizing the model itself instead of the data feeding it. Closing this gap requires observability at the feature level, not just the service level, so teams can understand what changed, when it changed, and why it mattered. Some teams pair this with isolated experimentation environments that allow them to validate changes safely before rolling them out broadly. 5. Scaling Real-Time Inference Reliably As models mature, they tend to depend on more features from more systems. What begins as a simple model using a handful of batch-computed features can evolve into a real-time service that requires multiple low-latency lookups across APIs, databases, and caches. While batch pipelines can scale relatively easily for training, real-time inference introduces strict latency and reliability constraints. Each additional dependency increases cost, complexity, and the risk of failure. Systems that worked well for experimentation often become fragile when exposed to production traffic. At this point, real-time inference stops being a modeling problem and becomes a systems problem: one that demands predictable latency, efficient execution, and careful control over what data is fetched and when. 6. Organizational Fragmentation Across Teams MLOps challenges are rarely purely technical. They are often amplified by organizational boundaries. Data science teams typically define features against warehouse tables, engineers reimplement those features against production systems, and MLOps teams manage infrastructure separately. When a schema changes in a warehouse or a key changes in a cache, a downstream model owned by another team may break without anyone realizing it immediately. Without a shared source of truth connecting data, features, and models, coordination becomes reactive rather than proactive. Changes are discovered through performance regressions rather than code review, and teams spend more time responding to incidents than improving systems. 7. Security, Access Control, and Compliance Constraints Production ML systems routinely touch sensitive data, including transaction records, identity information, and behavioral events. Teams must control access across environments, audit how predictions were made, and demonstrate compliance. When access controls, auditing, and governance are bolted on after the fact, they often slow teams down or block production deployments entirely. Security and compliance become obstacles rather than built-in properties of the system, making it harder to ship models confidently in regulated or high-stakes environments. How Chalk Can Help At their core, most MLOps challenges stem from fragmentation: different systems for training and serving, duplicated feature logic, and limited visibility into how models behave once they are live. Teams running high-stakes systems (such as real-time marketplaces, fraud detection, and credit decisioning) use Chalk to eliminate this fragmentation in practice. Chalk provides a single place to define, compute, and serve features using the same code across experimentation and production, batch and real time. Feature consistency is enforced by design rather than left to convention. Built-in lineage, versioning, and feature-level observability make it possible to trace predictions back to their inputs and understand what changed when model behavior shifts. Because Chalk runs directly in a team’s own cloud and integrates with any data source (including warehouses, operational databases, streaming systems, and APIs), teams do not need to move or duplicate their data. They get low-latency inference, predictable performance, and the ability to scale without maintaining parallel systems. The result is simpler infrastructure and more reliable machine learning. Instead of stitching together pipelines and hoping they stay in sync, teams can focus on building production ML systems that are easier to reason about, easier to maintain, and faster to ship. [See how teams use Chalk to run real-time ML] → Customer Stories [Read the docs] → Docs # Quarterly Product Update: Fall 25 | The Chalkboard source: https://chalk.ai/blog/product-update-nov-2025 This quarter, we expanded Chalk across three core dimensions: flexibility, visibility, and efficiency. Teams can now define and serve models alongside features, trace query behavior for diagnosis, and adapt feature logic dynamically without redeploying. These updates extend Chalk’s reach as a data platform for AI and ML, unifying how teams build, deploy, and observe systems for production and training. Here’s what’s new. Register and serve models alongside your features Chalk supports registering and running models directly in Python. Models are type-checked against their inputs and outputs, and referenced by stable names in Named Queries. This allows teams to track, version, and serve models in the same workflows they use for features. Model registration reduces infrastructure overhead and keeps training, batch, and live inference consistent by design. Read more about model registry in our docs. Reach out to our team if this is something you are interested in adopting. Tracing for query performance diagnosis Tracing gives teams deep visibility into how queries run inside Chalk. Each resolver and model call is instrumented and timed, making it easy to identify performance bottlenecks and understand why a query behaves the way it does. By surfacing this data directly in the dashboard, Chalk helps teams diagnose and optimize production systems with precision. Traces can be enabled per query through the CLI `--trace` or configured to sample automatically in production. This allows high-volume teams to monitor performance continuously without overhead. Vector aggregations for embedding workflows Chalk has improved support for vector search and implemented vector aggregations like `vector_sum` and `vector_mean`, enabling teams to compute over sequences of embedded features within Chalk. These capabilities power use cases like recommendation systems and fraud clustering, where teams need to perform time-decaying nearest neighbor retrieval or compute rolling representations of user behavior over time. By operating on vectors natively, teams can simplify architecture and reduce inference latency without leaving Chalk. @features class Transaction: id: int ... embedding: Vector[384] @features class User: id: int transactions: DataFrame[Transaction] ... mean_embedding: Windowed[Vector[384]] = windowed( "30d", "60d", "90d", materialization={"bucket_duration": "1d"}, # Vector mean will be computed as the element-wise # mean for each component of the vector expression=_.transactions[_.embedding].mean(), default=Vector(np.zeros(384, dtype=np.float16)), ) python These operations often feed model inputs for retrieval and ranking tasks and integrate cleanly with Chalk’s model registry and inference workflows. Dynamic Expressions for runtime logic Production feature sets increasingly require runtime variation without redeployment. Dynamic Expressions allow feature logic to be computed at query time instead of only at deploy time. With this capability, you can overlay context‑specific rules, support hot‑fixes, and run time‑aware logic without additional releases. We also expanded the expression library with: - Over 50 new Velox functions, including `max_by_n`, `min_by_n`, and aggregations like `approx_top_k` - Support for sklearn classifiers, now hostable directly in Chalk These updates make feature logic more expressive while maintaining type safety and performance. result, err := client.OnlineQueryBulk( t.Context(), chalk.OnlineQueryParams{}. WithInput("user.id", []int{1}). WithOutputs("user.name"). WithOutputs("user.email"). WithOutputExprs( expr.FunctionCall( "jaccard_similarity", expr.FunctionCall("lower", expr.Col("name")), expr.FunctionCall("lower", expr.Col("email")), ). As("user.name_email_sim"), ), ) Improved diffs and offline input visibility Production ML systems require strong introspection into what changed, what ran, and why. This quarter, we focused on surfacing that context more clearly across environments, deployments, and queries. - The improved diff viewer now renders full file contents when resolvers or features are added, removed, or modified. This makes it easier to review code changes, compare logic between branches, and trace regressions back to specific changes. - Dashboard metadata now includes runtime details like service account, namespace, and role ARN. This makes it easier to understand which services are running where, and under which policies. - Offline Inputs Explorer now shows the raw SQL and parameters used to build datasets. This supports validation of training workflows and better visibility into how feature inputs are sourced. Prompt templates for Cursor and Claude via CLI We added a new CLI workflow to help teams quickly generate structured prompt templates for LLM-assisted development. Running `chalk init agent-prompt `scaffolds prompt definitions that work with tools like Cursor or Claude and provide context on Chalk-specific syntax, including: - Creating feature classes and joins - Writing python and SQL resolvers - Adding unit and integration tests These templates help teams standardize development workflows and accelerate onboarding with built-in Chalk best practices. Templates are here. Chalk in production: recsys, credit decisioning, and underwriting World class teams continue to adopt Chalk across industries where real-time decisions matter. - Whatnot uses Chalk to power real-time recommendations across its live streaming marketplace, from feed ranking to show discovery. Read or watch their story here. - iwoca, one of Europe's leading small business lenders, adopted Chalk for its compute-first architecture to unify training and production pipelines, reduce drift, and make data consistency a system-level guarantee. Read their story here. - Mission Lane helps Americans build better financial futures by making credit more accessible. In a recent live demo, their Head of Data Science and ML showed how they manage thousands of features across training, live decisioning, and batch workflows, cutting rollout time from weeks to days. Watch it here. Learning & deep‑dive resources To help teams understand where and how Chalk fits into ML stacks, we published: - Chalk for data engineers: How to productionize features using Chalk’s Python interface without managing separate pipelines for training, serving, and batch workflows. - Chalk for data scientists: How to prototype features, run experiments, and deploy models with consistent features across training and production. - Chalk’s YouTube channel: Architecture deep‑dives, customer demos, and walkthroughs. # Mission Lane Demo Recap source: https://chalk.ai/blog/mission-lane-demo-recap How Mission Lane manages thousands of ML features in production using Chalk. Demo recap covering feature reusability, live decisioning, and batch workflows. We just wrapped up a live demo with Mission Lane, where Mike Kuhlen, who leads data science and ML, walked us through how they manage over thousands of features in production for over two million customers. Mike showed us real code, query plans in the Chalk UI, and explained how they handle complex feature dependencies across training, live decisioning, and batch evaluation. The setup Mission Lane helps people build better financial futures by making credit more accessible and transparent. Their ML models power credit decisions and fraud detection, pulling features from credit bureau data, internal transaction history, behavioral signals, and fraud providers — all building on each other in complex DAGs. This creates treacherous territory for feature updates. Before Chalk, shipping a new feature meant coordinating across data science, data engineering, and ML engineering teams. Even a minor change risked breaking downstream dependencies, creating weeks of delays and drift. (Sounds familiar? We've written about the common challenges in MLOps before). Demo walkthrough Define once, use everywhere Mike started by showing us their GitHub repo structure. Mission Lane defines each feature once, then reuses it across across real-time apps, monthly batch jobs and model training. He walked through an example with TransUnion credit bureau data. Even though they pull this data in different formats at different times, the feature definitions stay the same. Add a new bureau feature, and it automatically works everywhere without reimplementation. Decoupling models from infrastructure When a customer applies for a card, Mission Lane needs to score them in real-time. Instead of hardcoding all the features each model needs, they use named queries: request features by model name, and Chalk figures out which 300+ features to fetch. Whenever a model needs different features, just update the Chalk config — no need to touch the decisioning system code! Handling dependencies at scale Mission Lane scores millions of customers for credit line increases. This means processing hundreds of aggregated features with complex dependencies — aggregations depend on raw data, velocity features depend on aggregations, and so on. Mike showed us the query plan for one of their production models — a complex tree of hundreds of interdependent features that Chalk resolves together. Chalk's scheduled resolvers pull fresh data incrementally from Snowflake, tracking what's already been ingested and only fetching new records. By morning, the offline store is current and ready for batch scoring. (Chalk can also handle orchestration natively — some teams use it as a replacement for tools like Airflow and Dagster.) What stood out The real win: eliminating coordination overhead. Before Chalk, updating features meant orchestrating across multiple teams and systems, with changes rippling through complex dependency chains. Now data scientists self-serve with Python code that works everywhere, from real-time decisioning to batch evaluation! Watch the full demo to see the code, query plans, and orchestration: # Fast Company Next Big Things in Tech source: https://chalk.ai/blog/fast-company-2025 Chalk named to Fast Company's Next Big Things in Tech 2025 for helping AI teams get fresh data to their models in real-time. We're excited to share that Chalk has been named to Fast Company's Next Big Things in Tech 2025 list! The annual list spotlights emerging technologies with real potential to reshape industries over the next five years, from applied AI and infrastructure to sustainability and robotics. Fast Company's editors selected 137 honorees across 31 categories, looking for projects that represent genuinely new ideas and hold serious potential for growth. Their take on Chalk: we make "AI infrastructure a lighter lift." The Chalk approach Critical ML use cases, like identity verification and fraud detection, need fresh data in real-time to make accurate decisions. But the traditional approach forces a tradeoff: serve pre-computed data quickly (but stale) or fetch real-time data on-demand (fresh but slow). Chalk solves this by computing features at inference time, pulling fresh data directly from sources without ETL jobs. It's engineered for performance, built for ease of use. For example, data scientists write features in their notebooks in Python, and Chalk compiles them to optimized C++ automatically. Leading companies like Whatnot, MoneyLion and Socure rely on Chalk to power ML at scale. Earlier this year, we raised a $50 million Series A led by Felicis to continue building the future of real-time ML infrastructure. Building forward We're honored to be recognized alongside other teams like Runway, GitHub and Roblox tackling hard problems. Check out the complete Next Big Things in Tech 2025 list to see the full lineup. Here's to making data infrastructure for AI and ML faster, fresher, and a whole lot easier to build! # Chalk in Your Cloud | The Chalkboard source: https://chalk.ai/blog/deploy-in-your-cloud Chalk runs inside your VPC on Kubernetes so sensitive data stays resident, latency drops, and ML teams move faster. Hear the full story on the podcast. Enterprises building fraud, risk, recommendation, or search systems face three fundamental conditions: sensitive data must remain in-account, latency must be predictable, and control must stay with internal teams. That’s why Chalk is designed to run in your cloud, inside your VPC, on your Kubernetes. Feature computation executes adjacent to your systems, not across networks you don’t own. On Adventures in DevOps, Chalk co-founder Andrew Moreland describes why this approach is foundational to how we operate. Deploying inside customer accounts isn’t just an architectural preference, it’s a strategic differentiator that makes Chalk viable for enterprise workloads at scale. Why “in your cloud” wins: 1. Data stays in-account. Every component runs within your environment. Chalk can be configured for zero runtime access, so compliance, auditing, and governance remain entirely under your control. 1. Latency that scales. By co-locating feature computation with your applications, Chalk reduces cross-cloud hops and delivers ultra low latency for real-time decisions. 1. Correctness by design. Chalk’s data platform is inherently time-aware (temporal), preventing “future leakage” into training data and ensuring consistency across online and offline pipelines. How the advantage compounds: 1. No egress overhead. Compute and storage stay together, eliminating cross-account transfers and reducing interconnect costs. 1. One control plane across clouds. Kubernetes normalizes deployments. 1. Upgrades on your schedule. We ship versioned Helm charts and SDKs. You choose when to adopt. 1. Built-in moat at scale. Running on your cloud allows Chalk to sustain throughput, control p95s, and support complex feature logic without creating operational overhead. Listen to the full story. Andrew dives deeper into how Chalk deploys inside customer clouds, manages IAM cleanly, and ensures ML feature computation is both fast and correct. Watch the episode below: Learn more about how Chalk deploys in your cloud. # Chalk for Data Scientists source: https://chalk.ai/blog/chalk-for-data-scientists Your notebook code is production code with Chalk. Deploy directly from notebooks and experiment safely with branches. Let's talk about what modern data science actually looks like. CUDA, Scikit, and PyTorch were once nice-to-haves but are now table-stakes. These data tools you've mastered are now just the entry fee. Modern ML systems need real-time data pipelines, distributed computing, monitoring, and a dozen other things that probably weren't why you got excited about data science in the first place. Expecting data scientists to manage Kubernetes clusters and build observability systems is like asking a cardiologist to also perform heart surgeries: overlapping areas of expertise, but fundamentally different skill sets. This creates the bottleneck every data scientist knows: promising models sitting in notebooks, waiting for engineering resources that never come, ultimately never making it to production. Welcome to the Jupyter graveyard. But what if your notebook code was already production code? With Chalk, you can test new features, run experiments, and ship models to production all from your Jupyter notebook. @features class Review: id: int created_at: FeatureTime review_body: str rating: int @features(tags=["team:DevX"]) class Product: id: int title: str reviews: DataFrame[Review] average_rating: float | None = _.reviews[_.rating].mean() most_recent_review_rating: int | None = F.max_by( _.reviews[_.rating], sort=_.created_at, ) review_count_all: int = _.reviews.count() review_count: Windowed[int] = windowed( "30d", "90", expression=_.reviews[_.created_at > _.chalk_window].count(), ) python Why can't I just deploy my notebook? Hint: You can! Data Scientists traditionally would hand-off their notebooks to another team to be rewritten into another language (eg. Scala) for production. But with Chalk, the same Python you write for exploration runs in production with millisecond latency. We built a Symbolic Python Interpreter that converts ordinary Python into expressions that run natively. Easily import your production features with a single line of code, add a new experimental feature, and simulate a prediction, recommendation, etc. all from a single notebook! from chalk.client import ChalkClient client = ChalkClient() client.load_features() Review.is_positive = _.rating >= 4 Product.positive_reviews_percentage = _.reviews[_.is_positive == True].count() / _.review_count Experiment without breaking things Remember those week-old CSV dumps sitting in your folder? We've been there. You're trying to test a model that will run on live user data, but you're stuck with stale exports that don't capture what's actually happening in production. With Chalk branches, you can spin up a copy of your production pipeline in seconds, like Git but for ML features. Test against millions of production records, catch edge cases you'd never find in samples, and merge when ready. Iterate at the speed of thought, not infrastructure. chalk apply branch # deploys your model to a branch for testing chalk diff # prints a diff of changed features sh Once you're in a branch, changes hot-reload instantly: edit a feature definition, run a query, see results. Watch your new model predictions update in real-time as you tweak the logic. On Chalk, your features ARE your code, version-controlled and shareable just like any codebase. When colleagues need to test your features, whether it's another data scientist reproducing your work or a data engineer ensuring pipeline compatibility, they simply check out your branch. Train on real production data Here's how Chalk helps data scientists create accurate training datasets. Offline Queries: Every feature computed during production gets stored. Need to debug why your model approved a suspicious transaction last week? Query the exact features inputted and outputted at every stage of your feature pipeline. Temporal Consistency: Models often fail because they trained on future data that leaked into past features. Chalk handles point-in-time correctness automatically: OfflineQuery( name="Reviewed products from 2025 Q1", input={Product.id = list(range(10**5))}, output=[ Product.average_rating, Product.most_recent_review_rating, # these are computed within the context of the time window Jan 2025 - March 2025 ], recompute_features=True, lower_bound=datetime(2025, 1, 1), upper_bound=datetime(2025, 3, 31), ) Query features from March 15th, get exactly what existed on March 15th, without future transactions leaking into historical risk scores or manual timestamp juggling. Backfills Made Simple: Adding a new feature traditionally takes a week coordinating with data engineering. With Chalk, get it done in minutes. Change your feature definition, trigger a backfill. Chalk recomputes across your entire history while maintaining temporal consistency. Test "what if we had this feature 6 months ago?" before deploying. Ship it Once you've defined your features and resolvers, deployment is simple: chalk apply Your features now serve predictions at scale, and Python gets compiled to C++ for production performance. Integration with your existing ML stack is seamless. Deploy ONNX models directly in Chalk, or plug into SageMaker and Vertex: @features class User: id: str encoded_sagemaker_data: bytes should_offer_discount_percentage: float = F.sagemaker_predict( _.encoded_sagemaker_data, endpoint="discount_model-2025-01-15", target_model="model_v2.tar.gz", target_variant="blue", ) Plus, Chalk's native Iceberg integration enables sharing datasets with other teams, democratizing data access across the entire organization. Analysts can explore them in BI tools, everyone works from the same source of truth. # CEO Marc on The Information TV source: https://chalk.ai/blog/marc-on-the-information-tv CEO Marc on The Information TV unpacks why winning at AI isn't about bigger models but smarter infrastructure that keeps data and compute in your cloud. Our CEO Marc joined The Information TV to unpack why winning at AI isn't about bigger models — it's about smarter infrastructure. The infrastructure reality check Host Akash Pasricha cut straight to the chase — AI costs started falling, then suddenly plateaued. What gives? Marc explained that while chips keep getting more powerful, AI workloads are scaling even faster with larger models, more parameters, and bigger context windows. The demand is simply outpacing the hardware. His key insight: "Costs don't fall just because models improve. Costs fall when you get more efficient at what you're doing and when you learn how to use infrastructure in the right way." What's particularly interesting is that while everyone obsesses over training costs, the real money drain has shifted to inference — running these models again and again for actual users. Managing inference costs Marc shared practical approaches to managing costs, such as rightsizing your models (since not everything requires GPT-4 class reasoning) and caching smartly when you can tolerate values calculated in the past. But here's where Chalk shines: getting fresh data to models at inference time. He shared how real-time inference transforms e-commerce — new users get relevant suggestions from their first few clicks, not generic results from the previous night's batch job. As Marc put it: "Faster computation, on-demand computation, fresher data creates less fraud, more revenue, better user experiences." The data ownership debate Akash wrapped with a big question about the "corporate data wars" — who really owns the data when AI agents want to pull from every app in your stack? Marc's stance was clear: Enterprises need to own their own data. That's why Chalk deploys directly into customer environments where their data never leaves their cloud! Watch the full conversation There's plenty more we didn't cover, from why Chalk competes with in-house solutions (not Databricks), to how on-demand computation changes the cost equation. Watch the full episode on The Information TV to hear Marc explain why faster, fresher data is the difference between delighting users and losing them to competitors. P.S. Congrats to Marc on becoming a dad! His "no screens for a few years" policy might be the most ambitious goal he shared. # Chalk on TWiST500 source: https://chalk.ai/blog/chalk-on-twist500 Our CEO Marc Freed-Finnegan joins This Week in Startups (TWiST) to explain why AI needs fresh data fast. Learn how Chalk eliminates the speed vs. freshness tradeoff. "When things get faster, everything gets better." That's how our CEO Marc summed up Chalk's mission during his appearance on This Week in Startups (TWiST). As part of the TWiST 500, Marc sat down with host Alex Wilhelm to unpack why AI needs fresh data — fast. Speed as a feature Marc shared a lesson from his Google days: "Speed is a feature." When things are faster, people search more, click more ads, and everything becomes more magical. But companies today are forced to choose between fast data or fresh data. Traditional architectures either serve pre-computed data in milliseconds (fast but stale) or fetch real-time data on-demand (fresh but not instant). Take Whatnot, the live streaming marketplace. They were running overnight batch jobs for personalized recommendations — a standard approach that was burning compute on predictions that might never get used. With Chalk, they were able to both improve recommendations with real-time inference and save millions of dollars by deferring compute to when users actually need them. These are exactly the kind of problems Chalk solves — fetching fresh data directly from sources at inference time without the traditional latency penalty. It's how you get accuracy at a speed that feels magical. How we solved a recurring problem Marc revealed that co-founder Elliot built what would become Chalk three times — at Affirm, his own startup Haven Money, and Credit Karma. He would run into the same problem over and over again: Data scientists write models in Python, but it's too slow for production. Teams would rewrite everything in a faster language, introducing bugs and delays. After building an internal solution three times, Elliot realized how many teams were stuck in the same cycle. That's what sparked Chalk — why not solve this once for everyone? The approach is straightforward: we transpile Python to C++ automatically. Same Python your data scientists already write, but it runs fast enough for production without the manual rewrites. What’s next for Chalk Watch the full episode to hear Marc discuss open source plans, our SF/NYC/LA offices, and whether Databricks should be worried. If your ML models are stuck waiting for batch jobs, maybe it's time to see what happens when things get faster! # Chalk for Data Engineers | The Chalkboard source: https://chalk.ai/blog/chalk-for-data-engineers Data engineers waste weeks managing ML pipelines. Learn how to escape backfill nightmares and tool sprawl with declarative features on Chalk. Your daily Spark jobs run successfully. Your feature store is synced incrementally. Your operational dashboards report zero incidents with green across the board. Another team messages: “It seems like we’re not integrating refunds into our recommendation engine. We keep suggesting products that have high return rates.” Your stomach sinks because adding a refund filter requires recomputing six months of features and redeploying all of your ETL pipelines — in the correct order! Why ML pipelines feel like a chore Every new data source piles on complexity. Changing core business logic or even just integrating a new API requires updating multiple DAGs and rewriting the data contracts you have with the teams that depend on your feature logic. Real-time compute is riddled with bottlenecks. Warehouses take seconds for simple queries. Pre-computed features go stale quickly. Even worse, ML features need complex joins and aggregations — try computing "the likelihood a customer buys a product based on the similarity of their current browsing session with that of other customers" in 10ms. Any optimization made to squeeze out another millisecond risks breaking something else. This complexity is compounded by the fact that data engineers are drowning in tools: Airflow, Spark, Kafka, Flink, Redis — each of which is necessary for real-time. But stitching them together with glue code that is written in Python kills off any chance of being real-time in the first place. When something does break...inevitably, good luck debugging what went wrong! How Chalk simplifies everything Chalk abstracts away the complexity of doing real-time inference by enabling machine learning teams to express features and data relationships without worrying about data pipelines, caching, or serving infrastructure. Streamline feature engineering with declarative feature definitions and programmatic feature management: - Entity relationships that can be derived on the fly - Joining feature classes with composite keys that are computed in real-time - Built-in state management for incremental processing - Incrementally cache features with scheduled queries - Materialize (tiling) and window features as and with code - Persist individual features with multiple bucket durations at various cadences from chalk.features import DataFrame, feature, features, _ from chalk.streams import Windowed, windowed @features class Transaction: id: int created_at: datetime amount: float user_id: "User.id" user: "User" @features class User: id: int domain: str # composite keys that can be used as join keys workspace_id: str = _.domain + "-" + _.id expensive_api_call: str = feature(max_staleness="30d") # cache values # maintain different resolvers to A/B test function calls e.g. gemini vs openai llm_response: str = feature(version=3) # multi-attribute joins txns: DataFrame[Transaction] count_txns: Windowed[int] = windowed( "1d", "365d", expression=_.txns[_.created_at > _.chalk_window].count(), # https://docs.chalk.ai/docs/materialized_aggregations materialization=True, ) python One platform in your infrastructure Chalk runs in your VPC, inheriting your existing security groups and ACLs. Choose your own memory stores (Redis, DynamoDB) based on your performance needs. Your data never leaves your environment - full isolation and compliance. One less external service to manage, monitor, and secure. Any data source in minutes With Chalk, connecting a new data source is as easy as writing a new Python function or adding a .sql file to your Chalk repo. Remove cross-team dependencies, unblock data scientists, and iterate on your features with self-serve data access and A/B testing. -- resolves: Transaction -- type: online -- source: pg_txns -- incremental: -- mode: row -- lookback_period: 60m -- incremental_column: created_at select id, created_at, amount, user_id, from txns Smart caching and computation Define caching strategies down to the feature level. This prevents over-engineering while guaranteeing fresh data where it matters. - Automatic optimization: Set a staleness tolerance per feature — Chalk automatically re-uses features that have already been computed - Time-windowed aggregations: Compute feature aggregations at multiple intervals without building separate pipelines Build features, not pipelines Modern machine learning teams have been able to cut model development from months to days with Chalk. Stop spending a week setting up data pipelines just to test a new model idea. Connect new data sources, write features, and deploy a new model all in the same day. Chalk helps data teams focus on what matters most, creating value. Implement new features as fast as you can dream them up! # SciPy 2025 Recap | The Chalkboard source: https://chalk.ai/blog/scipy-2025-recap The Chalk team went to SciPy 2025 in Tacoma and came back inspired. Here's our recap of all the advances we saw in Python's data ecosystem. The Chalk team travelled up to Tacoma for SciPy 2025! Here’s what caught our attention. The composable Python stack has arrived Deepyaman Datta (a Kedro maintainer) demonstrated just how far the Python community has come! It’s now feasible to build a pure Python, composable data stack. Kedro + Ibis + dlt = Production pipelines in Python - dlt > Scaffolding for ETL pipelines with tens of thousands of sources and destinations - kedro > Data pipeline and orchestration framework - ibis > DataFrame library for that can execute on over 20 backends e.g. SQL engines With these three, you can extract and load massive datasets (dlt), orchestrate transformations and models (Kedro), and push computation down to the database or warehouse without leaving Python (Ibis). \# ibis t.filter(t.a > 1).select("a", "b", "c") python Turns into this SQL expression: SELECT a, b, c FROM t WHERE a > 1 sql Tag on DuckDB and you can build a lightweight data warehouse that runs locally. Pandera + Ibis = Data validation at warehouse scale - No more pulling data to memory for checks - Schema validation runs where your data lives GPUs go native Christopher Lamb from NVIDIA shared with the audience on how CUDA now has native Python support. No more C++ gatekeeping for GPU programming. The ecosystem now offers multiple paths to GPU acceleration. Through RAPIDS, cuDF provides true drop-in pandas acceleration, while Polars users get GPU power through the new cudf-polars backend. On that note: RAPIDS cuDF is making substantial progress integrating with Velox, an execution engine from Meta (that we optimize for online inference). Exciting times for those of us building high-performance data systems! The <10ms Challenge Our co-founder Elliot went down a different path, where fast inference is achieved through intelligent system design rather than raw compute power. His talk, "Real-time ML: Accelerating Python for inference at scale," tackled a specific challenge: How do you achieve sub-10ms latency while allowing teams to maintain the development velocity of Python? His key insights: - Parse Python's AST to build a computation DAG - Push filters and projections all the way down to the data source - Execute in Velox (Meta's unified execution engine) for massive parallelization - Functions that can't be transpiled (e.g. LLM calls) run in isolated processes with Ray Working with petabytes The Zarr team shared major updates for Icechunk, which is like Iceberg's scientific computing cousin; only solidifying the proliferation of separating storage from compute! Virtual references let you point to byte ranges in existing files without copying data - imagine building a unified view over petabytes without moving a single byte! Tom Nicholas from Earthmover demonstrated building massive datacubes from archival files. Instead of migrating petabytes of archival data to cloud-native formats, VirtualiZarr creates virtual layers that make existing files accessible as if they were cloud-optimized. The key insight: most scientific data is stuck in pre-cloud formats. Instead of a costly migration, create virtual layers that make old data look cloud-native. Notebooks and IDEs reimagined Unlike traditional notebooks, Marimo eliminates hidden state by treating notebooks as directed acyclic graphs of cells — making execution predictable, reproducible, and collaborative. It’s a batteries-included environment that can replace Jupyter, Streamlit, ipywidgets, Papermill, and more, with features such as: - Reactive and interactive workflows – run a cell and all dependent cells update automatically; bind sliders, tables, and plots to Python without callbacks. - Reproducible, shareable, and Git-friendly – deterministic execution, .py-based storage, built-in package management, and deployment as interactive apps or slides. - Easily export and re-remix your notebooks into a presentation, web app, or script! Together, these capabilities make Marimo a modern, end-to-end notebook environment built for both experimentation and production. Meanwhile, Posit (makers of RStudio) announced a new IDE built on VS Code's foundation but optimized for data workflows: - Integrated data viewers (not just print statements) - Variable explorers that understand dataframes - Native plot panes - Message: Data scientists deserve specialized tools Positron bridges the gap between lightweight notebooks and full IDEs, giving data scientists an environment that’s both powerful and purpose-built. It’s a reminder that general-purpose editors aren’t enough—data science thrives with tools designed for exploration, visualization, and analysis. Elvis has been using the editor since the announcement and has really enjoyed an AI-native notebook experience with the same powerful UX of VSCode. The art of going fast Brodie Vidrine from NOAA shared hard-won lessons migrating from Java to Polars in their Global Summary of the Month, scoring an impressive 80% performance improvement. When Polars wasn't enough for complex business logic, they dropped down to Numba for custom compiled functions. for v, v2, qc, qc2, dv in zip(tmin_vals, tmax_vals, min_qc_flags, max_qc_flags, daily_averages): daily_average_expressions.append( pl.when( pl.col("Element").is_in(["TAVG"]) & ~pl.col( v ).is_in([-9999.0]) & pl.col( qc ).is_in([""]) & ~pl.col( v2 ).is_in([-9999.0]) & pl.col( qc2 ).is_in([""]) ).then( ((pl.col(v) + pl.col(v2))/2)*.1 ).otherwise( -9999.0 ).alias(dv) ) Imagine dynamically building queries that compare over a hundred columns with just SQL! What this means for production ML SciPy 2025 marked a turning point. Python can now hold its own for enterprise workloads and has potential for becoming the primary user interface for interacting with production systems. Key trends include: - Unified execution: Write once, run anywhere (Chalk & Ibis) - Virtual data: Reference, don't copy (Icechunk) - Reactive development: Reproducible from the start (Marimo & Positron) The tools and ideas we saw in Tacoma weren’t incremental improvements — they’re signposts for where data and scientific computing are headed next. You can watch Elliot's talk here. # Quarterly Product Update: Summer | The Chalkboard source: https://chalk.ai/blog/product-update-august-2025 This quarter, we delivered platform-level improvements that extend Chalk’s ability to support complex, production-grade ML systems. This quarter, we delivered platform-level improvements that extend Chalk’s ability to support complex, production-grade ML systems. These changes focused on improving how teams define complex features, observe and debug runtime behavior, control infrastructure at scale, and integrate Chalk into their production ML workflows. Here’s what’s new. LLM feature support: evaluation and multimodal inputs LLM workflows are now first-class citizens in Chalk’s feature platform. This quarter, we launched native support for prompt evaluations, enabling teams to define, run, and compare prompts across models and datasets. import chalk.prompts as P chalk_client.create_dataset( input=df, dataset_name="Churn 2024", ) chalk_client.prompt_evaluation( dataset_name="Churn 2024", evaluators=["exact_match"], reference_output=User.churned, prompts=[ P.Prompt( model="gpt-4.1", messages=[ P.user_message("Determine the churn prob given login at {{User.last_login}}") ], temperature=0.2, ), P.Prompt( model="claude-sonnet-4-20250514", messages=[ P.user_message("Determine the churn prob given login at {{User.last_login}}") ], temperature=0.3, ) ], ) python Evals can also run directly from Chalk’s dashboard in addition to a Notebook. Metrics like latency, accuracy, and token usage are available alongside feature outputs. Once a prompt performs reliably, it can be deployed as a production feature, just like any other resolver. We also expanded the chat resolver, `chat.complete()`, to support multimodal prompts — combining image and text in a single call. This makes it easier to build GPT‑4o–style LLM workflows directly in Chalk. @features class Receipt: image_url: str image_response: P.MultimodalPromptResponse = P.completion( model="gpt-4o", messages=[ P.message( "system", [ {"type": "input_text", "text": "describe this image"}, {"type": "input_image", "image_url": _.image_url}, ], ), ], ) Feature flexibility Production systems need flexible building blocks. We added 50+ new Velox expressions, including support for scikit-learn (e.g. classifiers, regressions, random forest), data structure operations (`map_get`, `proto_serialize`, `struct_pack`), and various encoding functions (`sha`, `spooky_hash`). Scikit-learn models can also be hosted with Chalk instead of calling out to hosted model providers like Sagemaker. import chalk.functions as F from chalk import features @features class User: id: str total_spend: float avg_order_value: float days_since_last_purchase: int customer_tenure_days: int p_churn: float = F.sklearn_logistic_regression( _.total_spend, _.avg_order_value, _.days_since_last_purchase, _.customer_tenure_days, model_path=os.path.join( os.environ.get("TARGET_ROOT", "."), "churn_lr.joblib", ), ) SQL surface and acceleration visibility We improved support for SQL-native workflows and gave teams more transparency into performance: - We added native support for ClickHouse as a SQL data source. - We added the ability to easily query historical values (Offline Store), underlying data sources, and resolvers that return DataFrames. - We expanded our acceleration documentation to specify which Python constructs (e.g., `zip` and `enumerate`) are compiled into Velox expressions. Observability and runtime inspection Visibility into how features behave in production is critical. - We introduced workspace-level audit logs that capture API activity, user/service actions, IPs, and trace IDs. - We added support for Parquet exports of online queries, including inputs, outputs, query plans, and execution context. - We built an offline query input explorer to surface SQL and parameter inputs used during dataset builds. These additions improve auditability for teams running Chalk in regulated or high-risk domains. Infrastructure controls and SDK updates Data platform teams need tighter control over how compute is allocated and scaled: - We introduced per-pod rate limiting to control query throughput and concurrency under load. - We added node pool and pod-level isolation to enforce resource boundaries and prevent noisy neighbor effects. - We released a new TypeScript gRPC SDK with full support for `query`, `bulkQuery`, and `multiQuery`, matching the HTTP client interface. These changes reduce noisy neighbor issues and allow Chalk to behave predictably under bursty or latency-sensitive workloads. Chalk in action: lending, fraud, identity Real-world usage of Chalk continues to expand — here’s how customers are using it today: - MoneyLion uses Chalk for real-time fraud detection and personalization, accelerating time-to-production and improving feature freshness. Read more. - Mission Lane rebuilt its credit underwriting system with Chalk, moving to live features and explainable model inputs. Read more. - Verisoul uses Chalk for real-time identity risk scoring across heterogeneous data, blending ML into their decisioning system. Watch demo. Learning resources To help teams understand where and how Chalk fits into production ML stacks, we published: - Chalk for AI engineers: Learn how to build LLM and multimodal pipelines using prompt-time feature engineering. - Chalk’s platform architecture: Understand how Chalk is architected to support high-throughput, low-latency feature computation. - Evaluating features: Learn how to define, run, and analyze prompt evaluations at scale using Chalk’s native tools—combining dataset-backed metrics, model outputs, and deployment-ready workflows. - Chalk's YouTube channel: Explore technical walkthroughs, real-world demos, and architecture deep dives from our team and partners. # Verisoul Demo Recap | The Chalkboard source: https://chalk.ai/blog/verisoul-demo-recap Verisoul shows how to catch sophisticated fraudsters by using Chalk to combine ML features with LLM intelligence. See how they rapidly iterated through three different techniques in under an hour. We just wrapped up an inspiring live demo with Verisoul where Niel, their co-founder, walked us through how they catch fraud in real-time. Starting with 1,000 labeled domains, Niel built three different detection systems on Chalk, each more sophisticated than the last. This is exactly the type of challenge Verisoul tackles daily. They recently helped a major AI code editor discover they were bleeding $5M through free trial abuse — fraudsters were running 2,000+ free trials simultaneously from data center IPs! It’s the kind of brazen fraud that basic rules can’t fully catch. How Verisoul thinks about fraud Niel explained how Verisoul organizes their fraud detection into distinct modules — network intelligence, device fingerprinting, behavioral analysis. Their network module alone packs over 800 features, and that's just one piece of the system. These modules pull fraud signals from everywhere with database lookups, API calls, and even accelerometer readings. Before Chalk, stitching together those data sources meant weeks of engineering work. Now they can just define a feature, connect it to a data source with a few lines of code, and immediately use it across training and production. During the demo, Niel iterated through three different fraud detection approaches using real labeled data: 1. Basic domain intelligence Check if domains have valid email servers and how long they've been registered. Domains without MX records were marked suspicious, those under a month old were marked risky. The result: 42% accuracy. 2. Search result analysis Use an external search API (serper.dev) to fetch Google results for each domain, then scan for keywords like "disposable" and "temporary." This is because throwaway domains mainly appear in blocklists and security forums. Niel showed how simple this integration was in Chalk: just a Python resolver making a standard `requests.post()` call.* try: r = requests.post('https://google.serper.dev/search', json=payload, headers=headers) except requests.RequestException as exc: chalk_logger.error(f'Error fetching SERP results: {exc=}') python The result: 54% accuracy. *For production workloads, you can replace Python's `requests` with Chalk's native `F.http_get` expressions for faster HTTP calls! 3. LLM-powered analysis Here’s where things get interesting: use an LLM to generate a research report for each domain, and then another LLM to distill the report into a trust score. The result leaped to 96% accuracy. Seems like sometimes, the best feature engineering is letting an LLM do the pattern recognition for you! How Chalk makes ML + AI iteration fast The jump from 42% to 96% accuracy was impressive, but even crazier was how quickly Niel got there by testing three completely different approaches in under an hour. Before Chalk — what Niel called "the dark days" — each iteration would have taken weeks. So how did Niel iterate this quickly? A few things made it possible: 1. The code stayed simple. Every feature was just a Python function that would work in both development and production. 1. He tested everything on a branch server, experimenting with production data without any risk to the live system. When he applied changes with `chalk apply --branch`, he got a completely isolated environment. 1. He could backtest immediately against his labeled dataset without waiting for ETL jobs or copying data around. Just `client.offline_query()` and see results in seconds. 1. Traditional ML and AI features worked seamlessly together. When Niel's final approach needed to combine Spanner queries, web scraping APIs, and LLM calls, they all became features in the same pipeline. This combination — safe experimentation, instant feedback, and zero infrastructure overhead — is what let Niel try "wacky" ideas like multimodal LLM analysis. When ML and AI features combine this naturally, the best solution no longer feels too complex to build! Next steps and resources If you want to see Niel's complete live coding session, including the Q&A where he dives deeper into Verisoul's architecture and fraud patterns, the full recording is available here. # Feature store vs. Feature engine source: https://chalk.ai/blog/feature-store-vs-feature-engine Discover the critical differences between feature stores and feature engines. Learn when to use each approach and how feature engines enable real-time computation and rapid experimentation for ML teams. Feature stores have become widely adopted for solving training-serving consistency and enabling feature reuse. They work well for many use cases, but teams often hit challenges when they need real-time features, want to experiment quickly, or scale to more sophisticated ML applications. This is where feature engines come in. What is a feature store? A feature store is a centralized system that manages and serves machine learning features for models to make their predictions. It stores features for reuse across training and production, keeping models accurate and reliable. What is a feature engine? A feature engine is a computation platform that executes feature logic on-demand, with intelligent caching. It includes all the storage capabilities of a feature store, plus much more. Key capabilities that feature engines add: - Fresh data at inference time instead of only batch schedules - Computation for complex operations in real-time - Rapid experimentation with instant deployments - Unified infrastructure combining compute, storage, and monitoring - Self-service workflows for data scientists Feature store vs Feature engine - Features are data records in databases - Stores pre-computed feature values - ETL (Airflow/Spark) to offline store → more ETL to sync online store - Serves stored values without computation logic - Features get stale in between batch runs - Need data engineers to set up features and MLOps to productionize new features - Testing changes require full pipeline runs - Need to set up separate systems for observability Traditional Feature Store - Features are Python class attributes with typed definitions - Computes features on-demand at query time - Fetches from live data sources, caches as needed, no ETL needed - Query planner dynamically optimizes execution paths - Features are always fresh from source data - Data scientists can self-serve features and deploy directly with Python - Branch deployments enable isolated testing and iteration - Built-in monitoring, alerts, feature lineage, and versioning Chalk Feature Engine 1. Storage vs. Computation Feature stores are databases that serve pre-computed values. Feature engines are computation platforms that execute logic on-demand. When you query a feature store, it looks up a value. When you query a feature engine, it runs a function — traversing dependencies, fetching fresh data, executing transformations. It's like the difference between reading a saved file and running a program. 2. Static vs. Dynamic Feature stores reliably serve what external systems have computed for them, which works well if your feature definitions rarely change. However, feature stores don't inherently understand your features, they just store what external systems computed. You can't ask "how was this calculated?" or "what would happen if I changed this?". Feature engines understand the complete computational graph. The feature catalog lets you search and discover all available features, see their usage patterns, and understand dependencies. When debugging, you see the entire path from raw data to the final feature. 3. Pipeline-Dependent vs. Self-Service With feature stores, changing feature logic (eg. adding a data source) requires data engineers to update ETL pipelines, run backfills, and wait for data to populate—typically hours to days before the feature is usable. This setup makes experimentation especially demanding. Feature engines transform this workflow. Data scientists write features as Python classes and define logic using three types of resolvers: SQL, Python, and Chalk expressions for optimized operations like windowed aggregations. The engine handles all orchestration. Since features compute on-demand rather than through pipelines, there's no waiting for data to populate. Simply put: feature stores serve pre-computed values and need manual ETL work to sync their online stores. Feature engines compute on-demand and handle both offline and online serving automatically. Should I use a traditional feature store or feature engine? If you find yourself duplicating the same logic across different use cases as you add models, it's time to invest in feature infrastructure. It allows you to define features once and reuse them, rather than rebuilding for each model. Start with a feature store when: - Simple use case with predictable feature needs - Batch freshness meets your requirements Graduate to a feature engine when: - Use cases demand real-time features served at low-latency - Multiple teams need reusable features for development velocity - Data freshness directly impacts model performance - Infrastructure complexity is slowing speed of experimentation and innovation The future of feature engineering Looking ahead, more teams are moving from batch to real-time ML. Batch processing means features are only as fresh as the last pipeline run, limiting model performance for use cases like fraud detection or personalization, where recent user behavior matters. Developer experience is also improving. Traditional feature stores require coordination between data scientists and engineers to build pipelines and run backfills. Newer systems prioritize self-serve workflows where data scientists can define and deploy features themselves, directly in their notebooks. Curious how Chalk goes beyond traditional feature stores? Check out our product pages or book a demo. # Chalk for AI Engineers source: https://chalk.ai/blog/chalk-for-ai-engineers Learn how AI engineers use Chalk to eliminate stale context, reduce tooling complexity, and ship AI features that work consistently in production. You built a customer support bot that analyzes tickets and suggests responses. In your notebook, it understands context — pulling recent interactions, checking account status, and drafting helpful replies. Three weeks into production, it's suggesting refunds for customers who’ve already received them, and you can't tell which prompt version is running or what context the model has been given. So what led to that faulty suggestion in the first place? Why productionizing AI is so hard The context that your model needs lives in various systems: account status, refund history, previous support tickets. You sync it all to a feature store via overnight ETL jobs. When a customer writes in, your bot fetches these pre-computed features and calls an LLM, but the context is already stale. The freshness problem hits immediately. The refund processed this morning isn't in last night's batch. The policy that changed an hour ago isn't reflected in your bot’s responses. The tooling complexity makes it worse. To build this system, you're juggling: - ETL pipelines that sync your data to a vector database - An embedding pipeline with retry logic and rate limiting - Fetching features in real-time (but your context is still from last night's batch) - A prompt versioning system for testing and compliance - Cost tracking and monitoring Every component is a separate system with its own failure modes. When the AI makes decisions on stale data or breaks randomly, you're debugging across at least five different tools. Furthermore, since different model providers have different syntax for structured output (if they even support it in the first place), you end up writing adapters for each one. AI as infrastructure: Unified ML + LLM pipeline What if you could treat AI like any other infrastructure? Ultimately, that's what it is — most agents are just for loops with API calls. The engineering principles that make any system reliable apply here too. That’s exactly what Chalk does — it brings engineering discipline to AI systems. In Chalk, there's no distinction between 'ML' and 'AI' features: they're all just features in a unified platform. Whether it’s PostgreSQL data, API calls or LLM outputs, everything maps to feature classes that are versioned, typed, and testable like any other code. Instead of ETL jobs populating a feature store and making separate LLM calls, Chalk pulls fresh data and computes everything on demand when requests come in. These features are cached and immediately available and reusable across all your models. What you can do with Chalk Start where you are Chalk features work directly in Jupyter and Colab notebooks. When you're ready to move beyond experimentation, Chalk integrates with your existing vector database and infrastructure. The code you write for exploration is the same code that runs in production — no translation layer or rewriting for deployment! Define features once, use everywhere In Chalk, LLM outputs are features: computed once, cached, and reusable across teams and use cases. Your sentiment analysis becomes a feature. Your entity extraction becomes a feature. Any downstream model or service can use them. @features class Ticket: id: str user_id: User.id User: User status: AccountStatus = has_one(lambda: Ticket.user_id == AccountStatus.user_id) llm: P.PromptResponse = P.completion( model="gpt-4o-mini", messages=[ # Chalk injects features into the prompt # include features you want with Jinja templates P.message("user", USER_PROMPT) ], output_structure=SupportResponse, ) escalate_to_human: bool = F.json_value( _.llm.response, "$.escalate_to_human", # from structured output ) personalized_message: str = F.json_value( _.llm.response, "$.personalized_message", # from structured output ) python Real-time context injection Chalk pulls fresh data at inference time using Jinja templates. Your prompts reference features, and Chalk ensures those features are current when the LLM runs. prompt = """ Analyze the user's support request and account data to provide a personalized response. Most recent message: {{Ticket.user.last_message}} Account context: - Tier: {{Ticket.status.tier}} (VIP: {{Ticket.status.is_vip}}) - Total spend: ${{Ticket.status.total_spend}} - Recent tickets: {{Ticket.user.recent_ticket_count}} (unresolved: {{Ticket.user.open_ticket_count}}) - Average sentiment: {{Ticket.user.avg_sentiment}} Generate a SupportResponse with: 1. personalized_message: Acknowledge their status and address their concern 2. escalate_to_human: True if VIP, high-value (>$1000), or negative sentiment pattern """ Every decision is traceable. You see exactly what data the model received, when it was received, and how it influenced the output. Native vector search Whether you're finding similar support tickets or searching product catalogs, define embeddings alongside your data and pre-filter however you want, e.g. by category, user segment, price range. Search only the vectors that matter for your query. Prompt versioning & evaluation Prompts become versioned artifacts you can test, compare, and deploy. Run evaluations on historical data. Compare model performance on accuracy, latency, and cost. Your prompts are now engineered assets, with full audit trails and rollback capabilities. AI in production FAQs A few questions teams hit when an AI feature moves from notebook to production. Why do AI systems fail in production? The usual cause is stale or inconsistent context. Data synced by overnight batch jobs is already out of date by the time a request arrives, so a bot that worked in a notebook starts acting on values that no longer reflect reality. Spreading retrieval, embeddings, and prompts across separate tools adds failure points that are hard to trace. What is the difference between an ML feature and an AI feature? In a unified platform there is no real difference. Database values, API responses, and LLM outputs all become features: versioned, typed, and testable like code. They're also reusable across models and services. How do you give an LLM fresh context at inference time? Compute the values a prompt needs from live data at the moment of the request, rather than reading them from a batch that ran hours ago. When the prompt references features that resolve on demand, the model works from current data instead of a stale snapshot. You can trace exactly what it received. The simpler way ahead The right abstractions dissolve complexity. We built Chalk to unify the AI stack: everything maps to typed feature classes, LLM outputs become reusable features, and data stays fresh by default. Ship AI like you ship code! # What is MLOps? Definition, Challenges, and Best Practices source: https://chalk.ai/blog/what-is-mlops What is MLOps? Learn the definition, meaning, and benefits of MLOps, plus key components, challenges, and best practices for deploying and maintaining ML models. MLOps was nobody's job title in 2019. But over the past few years, the term's search volume has hockey-sticked as people realize that productionizing models is harder than it seems. Simply put, MLOps teams make machine learning a reality by bridging research and exploration with reliable production systems. Take a typical scenario: your data scientists build a recommendation model that crushes the baseline in testing—potentially worth $2M annually for a typical e-commerce site. Three months later, customers still see generic bestsellers. By the time it ships, multiple rewrites have broken the core logic. The revenue barely budges. Ultimately, MLOps is more than just deploying models. It's about getting four very different teams to work in harmony. Each speaks a different language and that's where things tend to break down. Here's how ML models typically go from idea to production: select Data Scientists identify what data they need and request pipelines from Data Engineers Data Engineers build ETL pipelines to move data from production systems to warehouses Data Scientists use that data to create features in Python notebooks and train models MLOps Engineers translate the Python features to production languages (Java/Scala/Go) and deploy Why ML teams get stuck This process creates multiple friction points: Features work differently in production: When MLOps engineers rewrite features in another language, small differences in implementation could lead to large drifts in model performance. Testing changes requires complex infrastructure: Most teams have staging environments, but they rarely match production. Testing a single feature change means redeploying entire pipelines. Problems span multiple systems: When production accuracy drops, the bug could be in the notebook, the data pipeline, the feature translation, or the deployment config. Each system is owned by a different team. Sequential handoffs create bottlenecks: Each team blocks the next. Data engineers wait for requirements. Scientists wait for data. MLOps waits for finalized models. Weeks and months pass. These problems compound with unstructured data. Need to parse receipts for expense categories? Extract sentiment from reviews? Now you need AI Engineers to build LLM pipelines. But AI Engineers who understand both ML and LLMs are rare. Most teams cobble together separate systems that don't talk to each other. By the time a model reaches production, it's unrecognizable. Multiple rewrites. Multiple languages. Multiple systems. When something breaks, nobody has the full picture. But what if features written in notebooks just worked in production? What if every team could contribute without rewrites or handoffs? How Chalk breaks down silos Chalk eliminates these handoffs by creating a single platform where all teams contribute without translation layers: The new workflow - everyone contributes to the same platform: - Data Engineers define data relationships at a higher level of abstraction - Chalk handles the pipeline complexity - Data Scientists write features in Python that run with ms latency in production - AI Engineers extract structured features from unstructured data using LLMs and embeddings - MLOps enables self-service deployments with built-in governance With Chalk: No rewrites. No handoffs. No disconnected systems. Here's how each team works with Chalk: Data Engineers: Declarative pipelines without the plumbing - Define data relationships that work across training and production - Configure caching, materialization, and time-windowed aggregations declaratively - Connect to existing SQL sources (Postgres, Snowflake, Athena) with resolvers - No manual ETL needed - Chalk handles query execution and optimization Data Scientists: From notebook to production in minutes - Write features in familiar Python on any notebook - Test on branches with historical production data - Push directly to production with `chalk apply` MLOps Engineers: Deploy and monitor without becoming the bottleneck - Enable self-service deployments for all teams - Monitor model performance and feature drift - Configure gradual rollouts and version control - Focus on deployment strategies, not infrastructure management AI Engineers: Turn unstructured data into ML features - Extract structured features from documents, receipts, reviews using LLMs - Systematically evaluate which prompts and models extract the best features - Deploy LLM-extracted features that compute alongside traditional features - No separate vector DBs, embedding pipelines, or prompt tools to manage What's next? The promise of MLOps was always about velocity — shipping ML value faster. Chalk delivers on that promise by aligning teams around a shared platform instead of forcing handoffs between incompatible systems. This blog post is just the beginning. In upcoming posts, we'll dive deep into how each team can leverage Chalk: - Chalk for Data Scientists: From notebook to production in minutes - Chalk for Data Engineers: Spin up data pipelines without the orchestration overhead - Chalk for AI Engineers: Building production LLM apps that scale - Chalk for MLOps Engineers: Deploying production-grade infra quickly, often, and reliably # Data Council 2025 source: https://chalk.ai/blog/data-council-2025 Insights from Data Council 2025 on the shift toward real-time inference, declarative data architectures, Python-first interfaces, and simplified data stacks—plus a look at the community shaping the future of data. I’m Elvis, developer advocate at Chalk. While attending Data Council last month and engaging in numerous conversations, talks, and demos, I observed several emerging trends: 1. Real-time is no longer a nice-to-have but a necessity 1. Code-first declarative architectures are replacing manual workflows and ad-hoc scripts 1. Python is emerging as the standard programming interface for data tools and platforms 1. Builders want a modern yet simple data stack Let's explore these key themes by looking at innovative startups and new products from the conference that show these trends in action. Real-time is no longer a nice-to-have but a necessity Real-time inference stands out as the most dominant trend due to the rise of AI-driven applications and growing user expectations for snappy, responsive product experiences. One of my favorite product launches from the week was MotherDuck's release of "Instant SQL". Instant SQL dynamically shows results as you type, transforming the traditional write-run-debug cycle into an interactive, real-time conversation with your data. Instantly preview and debug each component of your query—from CTEs to complex expressions—to quickly pinpoint issues and maintain an analytical flow state where your best thinking happens. More than just flashy UX, it works with any data source DuckDB can query—from CSVs to S3 parquet files—and is a strong signal that real-time interactivity is becoming table stakes for next-generation data platforms. At Chalk, we're witnessing this shift firsthand and expect that it will only accelerate as AI agents increasingly demand fresh context and customers come to expect hyperpersonalized interactions and recommendations. Code-first declarative architectures are replacing manual workflows and ad-hoc scripts The second shift is architectural: teams are moving away from drag-and-drop graphical user interfaces (GUIs) and toward code-first, declarative systems that bring the ergonomics of infrastructure-as-a-code (IaaC) to data management. Rill is a perfect example of this trend, using SQL as the primary language to build dashboards while streamlining the entire journey from data exploration to metrics visualization in a single tool. Declarative configurations in YAML or SQL are replacing point-and-click interfaces for everything from data transformation to dashboard creation. This approach enables teams to separate business logic from technical implementation, creating systems that are more maintainable, version-controlled, and reproducible. As data stacks continue to evolve, this code-first paradigm will increasingly become the foundation for building scalable, collaborative, and transparent data ecosystems. Python is emerging as the standard programming interface for data tools and platforms The third shift focuses on elevating developer experience across the data ecosystem. At its core, this shift addresses two key preferences of data practitioners: working with tools programmatically, and doing so in Python—a language not only loved by many but also supported by a rich ecosystem of data libraries. The widespread adoption of PyTorch, Keras, Polars, and Pandas underscores the popularity of Pythonic interfaces in the data science ecosystem. Bauplan, a programmable data lakehouse, embraces this Pythonic approach while innovating with concepts like functions over tables, Git-like versioning for data, and serverless execution aiming to make data workflows as seamless and reproducible as modern software systems. Branch-based data versioning, powered by Apache Iceberg, allows for safe testing and experimentation with transformations on production data without risking corruption. I get that Iceberg is all the rage, but it's important to view it as an implementation detail that serves to elevate developer experience and create delightful user interactions rather than being the ultimate goal in itself. Builders want a modern yet simple data stack Open source projects like Iceberg, Arrow, DataFusion, DuckDB, and Substrait have created a rich landscape of plug-and-play tools that work together seamlessly regardless of vendor or origin. And while this interoperability has created unprecedented flexibility, it also introduces a paradox of choice as teams now face an expanding universe of tools and integration patterns. That’s why a complementary movement toward simplicity is gaining steam—what Bauplan calls “the Simple Data Stack” and what Crunchy Data calls “the Post-Modern data stack”. What these efforts share is a focus on seamless, integrated experiences. The companies of tomorrow—especially in a market that is characterized by the evergrowing purchasing power of developers—are those that remove complexity with the right abstractions... not add to it. Crunchy Data's Iceberg extension for Postgres is a great example: - Iceberg is a powerful format but is only valuable when accessible through user-friendly tools and interfaces. - Crunchy Bridge enables seamless data movement from PostgreSQL to Iceberg with minimal configuration and a familiar syntax reducing what would typically be complex data pipeline setup to just two simple commands without sacrificing functionality. - The system supports all DML operations, preserves transaction boundaries, and handles advanced features like row filters and TOAST columns. The future belongs to flexible, unified approaches that are built for the 99% and meet developers and practitioners alike where they are, rather than forcing them to learn entirely new paradigms. Final thoughts: Community is infrastructure The most exciting part of Data Council wasn’t the tech. It was the people. Every talk, demo, and late-night side conversation reinforced a shared belief: the future of data belongs to builders who prioritize simplicity, speed, and developer empathy. At Chalk, we’re proud to be part of that movement—and even more excited to help power the teams building what’s next. If you’re working on real-time inference, feature platforms, or smarter ML infra—we’d love to talk. # Which LLM Wins at Nolan Trivia? source: https://chalk.ai/blog/chalk-prompt-evaluation We put Chalk’s native prompt evaluation framework to the test by benchmarking top models on a custom Christopher Nolan trivia dataset—showcasing how teams can evaluate and deploy LLM workflows with production-grade rigor. Large Language Models have revolutionized how engineering teams explore, interact with, and leverage their data. With just a single prompt, LLMs can operate as a variety of specialized models—from analyzing unstructured data and automating routine workflows to orchestrating complex agentic systems. Currently, LLM workflows usually involve shuffling data between multiple tools and services. Teams still struggle to confidently evaluate and deploy LLM-powered systems into production. Chalk is changing that. In this post, we’ll show how teams can use Chalk to accelerate prompt development and evaluation—by running them alongside feature engineering and model inference in a single unified platform. And to demonstrate these capabilities, we answer the burning question: “Which model would win at Christopher Nolan trivia?” Background Chalk is a comprehensive data platform for real-time inference at scale. Our goal is to help machine learning teams move quickly from ideation to deployment for any complex workload. With just a few lines of code, engineers can incorporate new data sources, compute new features via Python functions, and deploy their models. This same workflow extends to LLMs, making them first-class citizens in your ML stack. To support LLM workloads, Chalk provides a unified interface for inference, reducing boilerplate and increasing reliability. With prompt evaluation, developers can now test, trace, and deploy prompts directly alongside traditional ML pipelines. This makes it easier to launch evaluations across prompts, models, parameters, and criteria using your own historical or curated datasets—without switching contexts or tools. Creating a benchmark dataset For this walkthrough, our problem statement is: which LLM performs best at Christopher Nolan trivia? We’re using this example not just for fun—but to showcase a repeatable workflow: how to define an evaluation dataset, run prompt tests at scale, and promote the best-performing configuration to production. We define a trivia set using questions sourced from Wikipedia and IMDb’s Nolan trivia pages. To scope the task and reduce ambiguity, we use a multiple-choice format. While this introduces the risk of lucky (hallucinated) guesses, it offers an entry point for evaluating model performance. Here’s a sample trivia question that exhibits information retrieval and some reasoning: Which of the following themes has NOT been explored in a Christopher Nolan film? A. The nature of time B. Personal identity C. Religious salvation ← (Correct Answer) All 50 questions used for this evaluation can be found here. To upload this dataset to Chalk, we first define our data model. We’ll have 3 columns: `nolan_trivia.id`, `nolan_trivia.question`, and `nolan_trivia.actual_answer` that will be read directly from a dataset. In python, this will look like: from chalk.features import features @features class NolanTrivia: id: int question: str actual_answer: str python Then, we can create the dataset by sending the CSV data to Chalk. from chalk.client import ChalkClient import pandas as pd nolan_trivia_df = pd.read_csv("/path/to/nolan-trivia.csv") chalk_client = ChalkClient( client_id="", client_secret="", ) chalk_client.create_dataset( input=nolan_trivia_df, dataset_name="nolan_trivia", ) Defining a prompt In Chalk, prompts are defined just like features—using strongly typed Python functions, composable logic, and declarative interfaces. Here, we’re building a prompt to maximize accuracy on Nolan trivia. In Chalk, a “prompt” is a bundle of model choice, messages, and parameters. We define a system message to establish behavior, and dynamically populate the question using Jinja templating from our dataset. This setup allows us to reason about input/output behavior in isolation—a hallmark of Chalk’s local-first UX over a real-time compute graph. system_message = """You are a Christopher Nolan expert with comprehensive knowledge of his filmography, artistic techniques, personal history, and career achievements. Your purpose is to accurately answer trivia questions about Nolan with precision and confidence. When presented with questions, carefully analyze each option before selecting the most accurate answer. If you're uncertain about any details, reason through what you know about Nolan's career, filmmaking style, and public statements to make an educated assessment. Provide brief explanations for your answers when appropriate, citing relevant films, interviews, or known facts about Nolan. You should be able to identify incorrect options by recognizing inconsistencies with established facts about Nolan's life and work.""" user_message = """Please answer the following Christopher Nolan trivia question. Provide just the letter of the correct answer. Do not provide any additional context or reasoning. The only correct answers are A, B, or C. Question: {{nolan_trivia.question}} Answer: """ Testing the prompt Before scaling up, we need to validate that the prompt works correctly across different models and fine-tune parameters. Chalk makes this straightforward: we provide a unified interface allowing users to run completions and inspect detailed usage and runtime statistics directly in the notebook environment. Every Chalk completion returns a structured `PromptResponse` object that saves important metadata around the response. This means we can access token counts, latency metrics, and retry information as simple class properties, making it easy to track costs and performance during development. Let’s test our prompt on a single question in a notebook: from chalk import _, functions as F, prompts as P # Sample 1 question sample_question = nolan_trivia_df.sample(1).iloc[0].to_dict() # Define the completion feature NolanTrivia.completion: P.PromptResponse = P.completion( model="gpt-4o", messages=[ P.message(role="system", content=system_message), P.message(role="user", content=F.jinja(user_message)), ], temperature=0.2, max_tokens=1, ) # Extract response information as attributes NolanTrivia.response = _.completion.response NolanTrivia.total_tokens = _.completion.usage.total_tokens NolanTrivia.total_latency_s = _.completion.runtime_stats.total_latency # Define an accuracy evaluator function NolanTrivia.correct = _.response == _.actual_answer chalk_client.query( input=sample_question, output=[ NolanTrivia.id, NolanTrivia.question, NolanTrivia.actual_answer, NolanTrivia.response, NolanTrivia.total_tokens, NolanTrivia.total_latency_s, NolanTrivia.correct, ], ) Running this test gives us a complete picture of how our prompt performs: { "nolan_trivia.id": 38, "nolan_trivia.question": "Which of the following themes has NOT been explored in a Christopher Nolan film? A. The nature of time B. Personal identity C. Religious salvation", "nolan_trivia.actual_answer": "C", "nolan_trivia.response": "C", "nolan_trivia.total_tokens": 217, "nolan_trivia.total_latency_s": 0.5016413039998042, "nolan_trivia.correct": true } json Great! The model correctly identified (C) as the correct answer. Aside: While experimenting, you’ll often discover model-specific quirks. For example, we found that Gemini models add a trailing \n, which would break our exact-match evaluation. The fix here is simple: add a stop parameter. This is exactly why we test samples before scaling up. Now, we’re ready for a larger-scale evaluation across all of our target models! Launching an evaluation Chalk's Prompt Evaluation abstracts away the boilerplate of evaluation. You only specify the key components: - prompts: The list of Chalk prompts to use for the evaluation. - dataset_name: The name of the input Dataset to use for the evaluation. - reference_output: The name of the feature to use as the reference output for the evaluation. - evaluators: The list of evaluation functions to use for the evaluation. The functions could be a custom or Chalk defined-evaluators, e.g. “exact_match”. import chalk.prompts as P completion_kwargs = { "messages": [ P.Message(role="system", content=system_message), P.Message(role="user", content=user_message), ], "temperature": 0.2, "max_tokens": 1, } gemini_kwargs = { **completion_kwargs, "stop": ["\\n"], } chalk_client.prompt_evaluation( dataset_name="nolan_trivia", evaluators=["exact_match"], reference_output=NolanTrivia.actual_answer, prompts=[ P.Prompt(model="gpt-4.1", **completion_kwargs), P.Prompt(model="gpt-4.1-mini", **completion_kwargs), P.Prompt(model="gpt-4.1-nano", **completion_kwargs), P.Prompt(model="gpt-4o", **completion_kwargs), P.Prompt(model="gpt-4", **completion_kwargs), P.Prompt(model="gpt-4-turbo", **completion_kwargs), P.Prompt(model="gpt-3.5-turbo", **completion_kwargs), P.Prompt(model="claude-sonnet-4-20250514", **completion_kwargs), P.Prompt(model="claude-opus-4-20250514", **completion_kwargs), P.Prompt(model="claude-3-7-sonnet-20250219", **completion_kwargs), P.Prompt(model="claude-3-5-sonnet-20241022", **completion_kwargs), P.Prompt(model="claude-3-5-haiku-20241022", **completion_kwargs), P.Prompt(model="claude-3-haiku-20240307", **completion_kwargs), P.Prompt(model="gemini-2.0-flash", **gemini_kwargs), P.Prompt(model="gemini-2.0-flash-lite", **gemini_kwargs), P.Prompt(model="gemini-1.5-pro-002", **gemini_kwargs), P.Prompt(model="gemini-1.5-flash", **gemini_kwargs), ], ) The `prompt_evaluation` function abstracts away the boilerplate. It loads the dataset, uses the prompts to fill out all the appropriate completion calls with jinja templating, unfolds usage and runtime statistics in their own columns, and compares the response with reference output. The result of this function is reference to another Chalk Dataset that we could load into a notebook and perform some analysis, but the Chalk Dashboard also presents this information neatly. And the winner of the trivia challenge? Claude Sonnet 4—a true Nolanite. Deploying the winner Deploying the best-performing prompt is as simple as adding it to your feature model and running `chalk apply`. You now have a production-ready LLM feature running alongside your other features—backed by the same evaluation lineage and observability stack. from chalk.features import features @features class NolanTrivia: id: int question: str actual_answer: str llm_answer: P.PromptResponse = P.completion( model="claude-sonnet-4-20250514", messages=[ P.message(role="system", content=system_message), P.message(role="user", content=F.jinja(user_message)), ], temperature=0.2, max_tokens=1, ) Why native evaluation matters Prompt engineering today is often a manual process: iterating in notebooks, trying prompt variants, copying responses to spreadsheets, and judging outputs by eye. That might work for toy tasks—but in production, you need measurement, reproducibility, and iteration at scale. Chalk’s native prompt evaluation framework turns LLM development into a first-class engineering discipline. By treating prompts as structured, versioned artifacts and combining them with dataset-backed metrics, developers can: - A/B test prompts and models - Optimize for cost and latency - Trace evaluation lineage - Deploy with confidence In this demo, we used a single prompt template across providers, but there’s more dimensions to explore: per-model prompt tuning, parameter sweeps, embedded context, retrieval, and even tool-use to surf the web. Latency and cost are also key deployment considerations. Chalk surfaces these metrics directly during evaluations, giving teams the visibility to make informed tradeoffs—without additional logging infrastructure. At Chalk, we want ML teams to focus on innovating on their ideas, while we take care of details under the hood to optimize your compute workloads and surface insights. From feature computation to model inference to LLM prompt evaluation—we’ve built the tools to unify your stack, simplify your workflows, and move from experiment to production with speed and reliability. # Chalk Raises $50M Series A to Power AI Inference source: https://chalk.ai/blog/announcing-chalk-50-seriesa-funding I'm incredibly excited to share that Chalk has raised a $50 million Series A at a $500 million valuation, led by Felicis, with participation from Triatomic Capital and our existing investors General Catalyst, Unusual Ventures, and Xfund. This funding marks a huge milestone and a powerful validation of our vision for real-time AI infrastructure. I'm incredibly excited to share that Chalk has raised a $50 million Series A at a $500 million valuation, led by Felicis, with participation from Triatomic Capital and our existing investors General Catalyst, Unusual Ventures, and Xfund. We’re especially thrilled to have Aydin Senkut from Felicis joining our board. This funding marks a huge milestone and a powerful validation of our vision for real-time AI infrastructure. Reflecting on Our Journey When we started Chalk, we knew real-time inference was critical for fintech. Over the years, we've discovered that its importance extends far beyond fintech—to identity verification, fraud prevention, healthcare, and e-commerce. We've experienced remarkable growth—not just financially, but in the deep trust customers place in us to support their most critical applications. Companies like Doppel, Sunrun, Whatnot, Socure, Found, Medely, and iwoca rely on Chalk daily, continuously expanding their use cases. For instance, Doppel uses Chalk to rapidly detect threats in real-time, significantly enhancing security without compromising performance. Witnessing such customer-driven innovation has been incredibly rewarding. Looking Ahead AI compute is shifting rapidly from training to real-time inference, creating new demands for fresh data and complex computations at the exact moment decisions are made. Existing solutions have enabled large, complex training workflows and feature stores (low-latency caches of pre-processed data), but real-time inference remains underserved. Chalk uniquely addresses this critical gap by providing infrastructure designed explicitly for instantaneous, intelligent decisions. Our mission remains clear: to deliver intuitive, powerful data infrastructure that integrates seamlessly with developers' favorite tools. As applications involving large language models (LLMs) grow increasingly complex, our platform evolves to handle sophisticated data interactions more effectively. This Series A funding significantly accelerates our ability to build a fully general compute framework, enhancing integration and usability, allowing developers to effortlessly create transformative real-time AI applications. Our Team At Chalk, we hire exceptional people who strive to be among the very best in their fields. Exceptional software craftsmanship is our language—how we communicate our vision to the world. Our team shares a deep commitment to excellence, continuously challenging themselves to grow and to help build Chalk into one of the world's leading software companies. Elliot, Andy, and I bring extensive experience from Affirm, Palantir, Haven (acquired by Credit Karma), Google, and Index (acquired by Stripe), tackling large-scale data infrastructure challenges. Our shared journey has deeply influenced Chalk’s principles and continues to drive our growth. We’re rapidly expanding our teams in San Francisco and New York. If you're inspired by the future of AI infrastructure, check out our open roles here. Thank You Finally, a huge thank you to our incredible team, visionary customers, dedicated partners, and supportive investors. We’re grateful for your continued trust and collaboration. Here's to building an exciting future together! # April Events Recap source: https://chalk.ai/blog/events-recap-april-2025 Chalk's team recaps their April 2025 tour across major AI conferences, sharing what they learnt on real-time ML infrastructure and LLM integration. It’s been a busy month! The Chalk crew has gone on a cross-country tour (IRL and URL) to share the exciting updates our team is shipping. We spoke, demoed, did a roundtable and learned a lot. In case you missed it, here’s what we got up to! NexGen Banking Summit (NYC) Our founders and GTM team dove into the fintech ecosystem at NexGen Banking Summit in NYC! Elliot led a roundtable on one of fintech's biggest challenges: Most fraud models aren’t fast enough to prevent payment losses. He broke down how modern inference architectures are closing this gap, how we achieve sub-10ms decision-making, and when to build custom tooling versus buying off-the-shelf. Marc spoke on his time as product lead at Google Wallet and founding Index (now Stripe Terminal). Drawing a through-line from the early mobile payments revolution to today’s latency requirements, he shared how infrastructure decisions are critical for creating lasting competitive advantages in the space. Beyond fraud detection, an interesting shift we observed was how financial institutions are exploring ML to optimize customer experience and HR processing. Banks are now increasingly looking at their data infrastructure as a platform that can serve multiple use cases. VeloxCon at Meta HQ (Menlo Park) Nathan and Chase traveled to MPK to present an in-depth technical dive on our Symbolic Python Interpreter. Nathan pulled back the curtain on how we transpile Python into highly optimized Velox expressions at query-plan time, showcasing specific optimizations for reducing the performance bottlenecks during expression initialization. In his talk, Nathan revealed what we discovered while profiling a customer query with 10,000+ expressions: a full 50% of CPU time was being spent just on expression initialization. By caching expression equality patterns and implementing rewrite optimizations, we reduced this overhead to 17%, and pooling ExprSets across queries with the same shape brought it down even further. Beyond our own work, VeloxCon hummed with conversations on bridging traditional data engineering with modern ML demands. Many talks explored how to maximize performance across both CPUs and GPUs while maintaining developer accessibility in ecosystems like Spark and Presto. Developers we spoke to wrestled with the same core challenge: how do we make data infrastructure work better for ML without forcing developers to abandon the tools they already know and love? If you’re curious, Nathan’s full talk is now available on Youtube! Agents & GenAI Infrastructure & Tooling Summit (Virtual) Our co-founder Andy presented at the virtual Agents & GenAI Summit, demonstrating how to blend structured data with LLM outputs for fraud detection. Using GitHub star manipulation as our case study (a CMU paper estimates there are 4.5 million fake stars!), he showed how traditional graph algorithms like CopyCatch can detect obvious patterns but fall short with sophisticated fraud. Coordinated networks distributing malware through phishing repositories require more nuanced approaches. The demo illustrated how Chalk streamlines fraud detection by abstracting away complex orchestration between data sources. Teams can extract insights from wherever their data lives (Postgres, 3rd party APIs, vector stores) while we handle the parallelization, caching, and execution details. This frees engineers to focus on building and refining models rather than wrestling with infrastructure. The same pattern works equally well for financial fraud, content moderation, and anywhere else that needs both speed and precision. OptimizedAI Conference (ATL) & Data Council (SF) Elvis split his time between Atlanta and SF: first, he traveled to OptimizedAI to demo Chalk, share why we built our Symbolic Python Interpreter, and how we're thinking about Iceberg and the role it plays in separating storage from compute. Afterwards, he flew back to SF just in time for Data Council 2025! Open standards dominated discussions at both events. There’s plenty of interest in how Apache Arrow and Iceberg are enabling seamless integration into existing data ecosystems, but this newfound interoperability introduces an interesting paradox: more flexibility often means more complexity as teams navigate an expanding universe of tools. What connected these conversations was a growing emphasis on simplifying the modern data stack. We saw widespread validation for our Python-first philosophy that democratizes access by meeting practitioners where they are, using languages they already know. This reflects our core approach at Chalk – evolving familiar tools to meet modern needs rather than forcing teams to abandon what already works for them. What's Next? After a month of sharing our work at conferences across the country, it's clear that teams everywhere are wrestling with the same challenges we are: building AI-native ML systems that deliver both speed and simplicity. If you're interested in ML infrastructure and facing similar challenges around latency, developer experience, or bridging structured and unstructured data, we'd love to connect! # Quarterly Product Update: Spring source: https://chalk.ai/blog/product-update-april-2025 This quarter’s updates help ML teams move faster in production with better performance, observability, and flexibility. Over the last quarter, we’ve shipped major upgrades to Chalk that make it easier for ML and data teams to ship features faster, iterate confidently, and power production models with real-time data at scale. We expanded support for Python acceleration, improved system observability, and gave teams more control over how features are persisted, joined, and queried. Write Python, Execute in C++ We expanded our Symbolic Python Interpreter to support more resolver types, enabling Chalk to compile more Python logic directly into C++ implemented Velox expressions. This gives you the ergonomics of Python with the performance characteristics of native execution. For teams deploying high-throughput or latency-sensitive models, this means: - Lower latency (sub-10ms) through native execution paths for more resolvers - Higher scalability through vectorized operations (SIMD & no GIL) - Faster development cycles with streamlined feature engineering (Python syntax—no DSLs) Learn how our engineering team built this, and watch Nathan present it at Veloxcon '25. Greater control over how data is modeled, persisted, and reused We introduced new patterns that give teams finer-grained control over how features are defined and managed in production: - Selective feature persistence: Exclude intermediate or expensive fields (like raw text or embeddings) from online and offline stores with `store_online=False` and `store_offline=False`. - New aggregations: Functions like `array_median`, `max_by_n`, and `array_sum` unlock new statistical modeling patterns out of the box. - Expressive filtering & joins: Use Chalk Expressions for `.where()` filters across DataFrames and has-many relationships. Support for composite key joins makes entity resolution in complex datasets much easier. @features class User: id: str = _.alias + "-" + _.org + _.domain org_domain: str = _.org + _.domain org: str domain: str alias: str # join with composite key posts: DataFrame[Posts] = has_many(lambda: User.id == Post.email) # multi-feature join org_profile: Profile = has_one(lambda: (User.alias == Profile.email) & (User.org == Profile.org)) @features class Workspace: id: str # join with child-class's composite key users: DataFrame[Users] = has_many(lambda: Workspace.id == User.org_domain) python Composite keys enable teams to model complex relationships that require multiple attributes for unique identification. This supports more accurate data modeling, flexible cross-dimensional joins, and proper data isolation in multi-tenant environments. New observability tools for real-time systems We made Chalk’s dashboard a more powerful tool for understanding system behavior, debugging edge cases, and optimizing for scale. - The Online Query Performance tab surfaces SQL queries, resolver inputs/outputs, and acceleration insights—giving full visibility into how queries are executed in production. - The Query Plan Viewer now highlights accelerated Python resolvers as yellow and displays the resulting static expressions they’ve been converted into. - Feature-level observability includes most-used values, value counts, and slice-based metrics for each deployment—helping teams debug drift and detect regressions quickly. - A new graph view for feature lineage makes it easier to understand how features are constructed, reused, and deployed across your system. These improvements are especially valuable for teams operating in real-time domains like fraud, credit, personalization, and infrastructure optimization—where explainability and performance tuning are critical. Customer spotlight: Apartment List and Verisoul Teams are putting these upgrades to work already: Apartment List rebuilt their recommendation engine with Chalk, delivering personalized results in real time by dynamically flexing price and location preferences and making low-latency API calls during inference. They went from nightly batch to always-fresh results—without refactoring their model. Read their story here. Verisoul uses Chalk to iterate on fraud detection logic with minimal friction. They ship new features in hours (not weeks), improving precision and recall while reducing operational overhead. Read their story here. Other highlights - Improved Glue Catalog performance: Distribute training datasets to downstream teams (analytics, data science, etc.) by writing offline queries to Iceberg and your catalogs. - Autoscaling with KEDA: Dynamically scale workloads based on event traffic (e.g. Kafka), now live across all customer clusters - Vertex AI embeddings support: Use GCP’s Vertex models with Chalk’s built-in `embed()` function. class AnalyzedReceiptStruct(BaseModel): expense_category: ExpenseCategoryEnum business_expense: bool loyalty_program: str return_policy: int @features class Transaction: merchant_id: Merchant.id merchant: Merchant receipt: Receipt llm: P.PromptResponse = P.completion( model="gpt-4o-mini-2024-07-18", messages=[P.message( role="user", content=F.jinja( """Analyze the following receipt: Line items: {{Transaction.receipt.line_items}} Merchant: {{Transaction.merchant.name}} {{Transaction.merchant.description}}""" ))], output_structure=AnalyzedReceiptStruct, ) vector: Vector = embed( input=lambda: Transaction.receipt.description, provider="vertexai", # openai model="text-embedding-005", # text-embedding-3-small ) We're continuing to invest in performance, observability, and developer experience—so teams can go from idea to production without slowing down. If you're building real-time ML and want to move faster, we'd love to show you what Chalk can do. Explore the full changelog. # Symbolic Python Interpreter source: https://chalk.ai/blog/symbolic-python-interpreter To eliminate Python’s performance bottlenecks in real-time ML, Chalk built a Symbolic Python Interpreter that converts flexible Python resolver code into optimized Velox-native expressions. This post dives into the technical journey and impact of enabling fast, parallel execution without sacrificing developer ergonomics. TL;DR Python is the language of choice for many ML teams — it’s flexible, expressive, and has the right ergonomics for real-time workflows. But it’s also slow. To make Python faster, I built a Symbolic Python Interpreter to make Chalk’s “Python resolvers” into optimized Velox-native expressions — enabling high-performance execution without sacrificing developer experience. This post explains the problem I encountered, why it’s so challenging, and how we solved for it. Background Machine learning teams want to move fast: ship models, incorporate new data sources, and react to real-time signals. But traditional feature pipelines — including ones written in Python — are slow, brittle, and not built for low-latency environments. At Chalk, we’re building a real-time feature platform to change that. Chalk is a real-time feature platform designed to help teams incorporate new data sources, and iterate on ML models faster for real-time decision-making. At its core, Chalk allows engineers to define and compute features dynamically, ensuring that models have access to the freshest and most relevant data. Chalk’s computation engine revolves around resolvers — composable blocks of computation that describe how data is stored or derived from existing data. Resolvers pull data from various sources like SQL databases or compute new values through Python functions, making feature engineering faster, more flexible, and easier. Python resolvers: flexible, but not always fast The most flexible type of resolver in Chalk is a Python resolver. With a Python resolver, you can: - Specify input features. - Write Python code to compute new features. - Use any Python library, including Pandas, Polars, NumPy, and even web-scraping tools like BeautifulSoup. Python is great for writing custom logic, but it comes with performance trade-offs. For example, consider a transaction-processing resolver that recalculates the expected total based on discounts, tax rates, and whether the transaction is taxable: def compute_total( subtotal: Transaction.subtotal, discount: Transaction.discount, tax_rate: Transaction.locale.sales_tax_rate, is_taxable: Transaction.taxable, ) -> Transaction.total: if discount is not None: # Apply any discount, but don't allow it to become negative. subtotal = max(subtotal - discount, 0.0) if not is_taxable or tax_rate is None: return subtotal return subtotal * (1.0 + tax_rate) python The flexibility of Python makes it easy to implement custom business logic, but running Python resolvers in real-time ML pipelines introduces performance bottlenecks. How Chalk executes Python resolvers When executing a query, Chalk’s query planner determines which resolvers need to run to compute missing features. Resolvers may: - Retrieve data directly from SQL databases (Postgres, Snowflake, BigQuery, etc.). - Compute values using Python. - Call hosted models (in Sagemaker, Vertex) in addition to 3P models and LLMs, and third-party APIs or perform additional transformations. Velox Execution Engine Once the query planner identifies dependencies, it constructs a Velox execution plan. Velox is an open-source execution engine that processes streams of table fragments efficiently. Chalk integrates Python resolvers by wrapping them as Velox operators, enabling execution at scale. Python scalar resolvers execute in dedicated subprocesses: 1. Data batches are dispatched to a Python subprocess. 1. The subprocess converts table data into native Python values. 1. The user-defined function executes on each row of data. 1. The results are reassembled into a Velox table. This design allows for the full flexibility of Python to compute arbitrary features, but introduces some performance challenges. Why is running Python slow? Python resolvers allow for powerful, expressive logic, but they come with trade-offs: - Global interpreter lock (GIL) - Within each process, only a single pure Python thread can be executed, so a single process cannot fully utilize all CPU cores when running pure Python. - Dynamic typing overhead - The Python interpreter must handle unpredictable types at runtime, introducing inefficiencies. - Heavy value representation - Each Python value (including `int/float`) is a heap-allocated, reference-counted object, which is memory-intensive. - Single-row processing - Modern CPU architectures include single-instruction-multiple-data (SIMD) operations which allow individual threads to perform many arithmetic operations in parallel. But since scalar resolvers run Python on each row of data individually (which provides the best developer experience), vectorization is not possible for queries with many rows. To mitigate some of these issues, Chalk runs Python resolvers in multiple subprocesses (potentially across many computers). However, this still isn’t enough to efficiently serve high query volume at scale. The origin of the Symbolic Python Interpreter Chalk already runs all of our built-in compute-intensive operations as Velox operations, which avoids all of the above issues — Velox is natively multi-threaded, uses efficient, statically-known layouts for all values, and performs vectorized SIMD operations for compute bulk arithmetic operations out-of-the-box. I built the Symbolic Python Interpreter to convert pure Python resolvers into equivalent Velox-native expressions at query-plan time, before execution. This allows our customers to write simple, easy-to-understand Python functions which operate on individual rows of data, but are executed as (potentially SIMD) operations on bulk tables with known type layouts. This transformation is conceptually similar to Python compilation tools like Numba, but Chalk’s compiler is specifically designed to support real-time ML workloads. How the Symbolic Python Interpreter works At query-plan time, each row-for-row Python resolver is scanned as a candidate for conversion into a Velox expression. The query planner obtains the resolver’s input types from its Feature Types class annotations – these are the types of the operator’s input table’s columns. The Symbolic Python Interpreter then executes the function symbolically at query-plan time. Instead of being provided specific concrete values, the function is called with symbolic values as inputs - these represent trees of computations (consisting of function calls, input columns, conditional branches, and constants). These symbolic values propagate through the function as it is executed symbolically. For example, a very basic function like @online def compute_total( subtotal: Purchase.subtotal, tax: Purchase.tax, ) -> Purchase.total: if tax is None: return subtotal return subtotal + tax becomes a symbolic tree of columnar operations: if_else( is_null(column("Purchase.tax")), column("Purchase.subtotal"), add(column("Purchase.subtotal"), column("Purchase.tax")), ) Each sub-expression within the tree tracks both the exact Python type (allowing Python’s dynamic semantics to be precisely emulated) along with the Velox type in the resulting expression (allowing the expressions to use efficient, compact layouts for each sub-expression). The interpreter also tracks control-flow information, so that flow-sensitive information can be used to eliminate dynamic checks and accelerate dynamic functions. For example, in the above example, even though `tax` has Python type `Optional[float]`, the interpreter is able to statically reason that `tax` must not be `None` in the final expression because the function would have already returned. Unknown or unsupported functions cause the interpreter to bail out and fall back to the original generic Python subprocess invoker, which is capable of running any Python code. At the end of the function, if symbolic execution was successful, the query planner has obtained a single complex Velox expression representing all of the Python resolver’s code, which emulates Python semantics exactly (including nullability, short-circuiting, and loops over list-typed features). Then, Python execution can be replaced entirely by a Velox projection operator which computes the resolver’s output as an additional column. How symbolic interpretation accelerates execution By bypassing Python’s runtime overhead, this approach delivers substantial performance improvements. The resulting expression uses a compact, efficient representation for each input, output, and intermediate sub-expression, avoiding the need for excessive heap allocations or dynamic checks on every function call and arithmetic operation. Velox natively runs all operators on multiple threads, allowing for trivial CPU scaling as multiple requests are handled concurrently or one request scales to very large numbers of rows. In addition, Velox operates on batches of rows, so each operation can trivially make use of SIMD instructions when dealing with multi-row inputs since the row-by-row function is converted into a columnar vector expression. The impact By eliminating Python’s inefficiencies, Chalk’s Python resolvers now: - Run in parallel without unnecessary overhead. - Translate to Velox-native expressions for optimized execution. - Execute faster without requiring users to write specialized performance-tuned code. Since Python resolvers are automatically accelerated, our users can focus on defining features and experimenting, instead of spending time on a tedious and extraneous “productionization” step. Chalk’s users can just write ordinary Python code, and it’s already ready to ship for production workloads. Final thoughts Our work on the Symbolic Python Interpreter changes how Chalk executes Python resolvers. By eliminating Python’s runtime inefficiencies, Chalk now: - Runs Python resolvers in parallel. - Identifies vectorizable expressions for Velox-native execution. - Uses a Symbolic Python Interpreter to analyze and optimize function logic automatically. By combining Python’s ease of use with a highly optimized execution engine, Chalk ensures that real-time ML workloads remain both flexible and fast. Stay tuned for further optimizations as we continue pushing the boundaries of real-time feature computation. Defining new feature resolvers used to mean choosing between expressiveness and production-ready efficiency. With Chalk’s accelerated Python resolvers, now you get both — flexible and fast. # Quarterly Product Update: Winter source: https://chalk.ai/blog/product-update-dec-2024 A roundup of Chalk's latest product releases, feature updates, and new documentation for the end of the year 2024. As the year comes to a close, the Chalk team has been hard at work delivering powerful new features to enhance your workflows, improve observability, and streamline integrations. This release is packed with updates designed to help you unlock new efficiencies, achieve better performance, and close out the year on a high note. As always, you can stay updated with our weekly changelog. More functionality for expressions Chalk expressions have been significantly expanded to simplify feature engineering workflows and boost performance through static analysis and C++ compilation. You can now reference fields within structs, and the expanded `chalk.functions` library enables creating logical expressions, transforming features, and integrating predictions directly into workflows by running SageMaker predictions with F.sagemaker_predict. Check out the full library of `chalk.functions` to unlock new transformations with less code and more speed. Dashboard improvements featuring more metrics and observability Several updates have been made to the dashboard to enhance observability across your features, resolvers, queries, and cluster details. The revamped overview page now provides comprehensive metrics, including insights into recent online and offline queries, resolver runs, deployments, errors, and the status of all connections in your environment. To enhance monitoring, we've added warning banners in the dashboard to alert you to issues in deployments, such as pods failing to start up cleanly. For deeper visibility, the deployments page now displays the kubernetes pod resources associated with each deployment. Additionally, for more granular observability, you can view the latest stack trace for each pod in your cluster and filter logs by pod name, resource group, deployments, and more. Enhanced filtering and exploration capabilities now make it easier to gain precise insights across the dashboard. In the resolver table, you can view p50, p75, p95, and p99 latency statistics, while customizing column selection for comparisons. The usage dashboard supports filtering and grouping by cluster, environment, namespace, and service. In environments with a gRPC server, the features page presents feature value metrics and recently computed feature values loaded from the offline store. Lastly, a SQL explorer has been added to the dashboard, allowing you to run SQL queries against datasets, including query outputs, for faster data exploration and analysis. Offline query improvements Offline queries are now more powerful and flexible with several key improvements. They can accept parquet files as input by passing in a `s3://` or `gs://` URL to the `offline_query(input="...")` parameter. You can also control the parallelization of offline query execution by specifying the num_shards and num_workers parameters. Additionally, upper and lower bounds for offline queries now support timedelta inputs, enabling dynamic time-based constraints, such as `offline_query(upper_bound=timedelta(days=30)` to set the upper bound to be 30 days after the latest `input_time`. Lastly, the store_online and store_offline parameters allow you to seamlessly store offline query outputs online and offline, respectively, for improved integration. Optionally cache null or default feature values in online store The `@features` decorator now includes cache_nulls and cache_defaults as parameters, which specify whether to cache and update null or default computed feature values in the online store. For DynamoDB or Redis online stores, these parameters also accept `cache_nulls="evict_nulls"` or `cache_defaults="evict_defaults"` to evict null or default feature values from the online store, prompting online resolvers to rerun and compute up-to-date results. Integration testing with ChalkClient The ChalkClient now supports integration testing for your features and resolvers through the `.check()` method. This method allows you to deploy local changes, query against branches in a codified way, and run integration tests in your CI/CD pipeline. from chalk.client import ChalkClient from chalk.features import DataFrame, features, FeatureTime, _ from chalk.streams import Windowed, windowed import datetime as dt import pytest @pytest.fixture(scope="session") def client(): return ChalkClient(branch=True) # Uses your current git branch @features class Transaction: id: int user_id: "User.id" amount: float ts: FeatureTime @features class User: id: int transactions: DataFrame[Transaction] transaction_count: Windowed[int] = windowed( "1d", "3d", "7d", expression=_.transactions[_.ts > _.chalk_window].count(), ) def test_transaction_aggregations(client): now = dt.datetime.now() result = client.check( input={ User.id: 1, User.transactions: DataFrame([ Transaction(id=1, amount=10, ts=now - dt.timedelta(days=1)), Transaction(id=2, amount=20, ts=now - dt.timedelta(days=2)), Transaction(id=6, amount=60, ts=now - dt.timedelta(days=6)), ]) }, assertions={ User.transaction_count["1d"]: 1, User.transaction_count["3d"]: 2, User.transaction_count["7d"]: 3, } ) python Miscellaneous features and improvements - A new chalk usage command has been added to retrieve and export Chalk usage data. - The chalk healthcheck command allows you to check the health of the Chalk API server and its services. - Poetry is now supported for managing Python dependencies, providing a streamlined way to configure your Chalk environment. Read more about how to configure your Chalk environment here. - Pub/Sub is now supported as a streaming source. Read more about how to configure your Pub/Sub source here. - An idempotency key is now available to ensure that only one job is triggered per idempotency key, preventing duplicate executions. # Quarterly Product Update: Fall source: https://chalk.ai/blog/product-update-september-2024 A roundup of Fall 2024’s latest product releases, feature updates, and new documentation and tutorials. It's back-to-school season and we're back with another roundup of what the Chalk team worked on for the past few months! As always, you can follow along for weekly updates at our changelog. Updated UI for features and resolvers The Features and Resolvers sections of the Chalk dashboard have new functionality allowing you to filter, sort, and resize columns! Additionally, the Features section shows the most recent request counts for each feature and has a button for downloading the table as a CSV report. Queries can reference multiple feature classes Previously, you could only reference one feature class in each query, which meant executing several queries to collect all of your features and then stitching the corresponding data back together. Now, you can request features from multiple feature namespaces in a single query! For example, here’s a query where we retrieve information about a specific user and how they interacted with a video: client.query( input={ User.id: "12345", Video.id: "98765", UserVideo.id: "12345_98765", }, output=[User, Video, UserVideo], ) python Named queries You can now define named queries, which allow you to document key use cases in code and execute queries just by referencing their name. Here's an example where we define a `NamedQuery` similar to the previous example with users and videos: from chalk import NamedQuery from src.feature_sets import User, Video, UserVideo NamedQuery( name="user_video_likes", input=[ User.id, Video.id, ], output=[ User.id, Video.id, Video.num_views_30d, UserVideo.id, UserVideo.watch_duration, UserVideo.interaction, ], tags=["team:analytics"], owner="analytics@example.com", ) After defining this query and running `chalk apply`, you can execute the query by name without enumerating all of the outputs: chalk query --in user.id=123 --in video.id=456 --user_video.id=123_456 --query-name user_video_likes As shown in the section above regarding our updated UI, named queries appear in the feature table in the "input to" and "output of" columns. More functionality for expressions Expressions have received several changes over the past few months! Expressions are a terse syntax for expressing simple transformations and aggregations of other features in the same feature class. - You can now use `_.chalk_window` in windowed aggregations to reference the current target window. - You can now use the `chalk.functions` module with expressions for common feature transformations. You'll find functions for unzipping, hashing, coalescing, and converting data types. Native support for PartiQL with DynamoDB We now support using PartiQL with DynamoDB! PartiQL is a great tool for DynamoDB because its queries can be translated into fast, native DynamoDB queries. As a bonus, our teammate Bill Qin wrote a deep dive on how he changed our PartiQL parser to support column aliasing for the sake of creating a better developer experience. It's always fun to see different use cases for manipulating ASTs in production! New tutorial for using Chalk with AWS SageMaker We have a new tutorial showing how you can use Chalk within a SageMaker pipeline for model training and evaluation. In short, Chalk is a great feature engineering platform because you can use the same code across your notebooks, training, and serving. Meanwhile, SageMaker shines for model training and serving. With their powers combined, you can build a streamlined model training and serving system that uses the best of both systems. Updated documentation - We updated our time documentation. Don't worry, the concept of time has not changed, but we hope our new docs are easier to digest. - We also updated the documentation for expressions. We've also added an index of all available expressions. # Chalk wins at GenAI Demo Day source: https://chalk.ai/blog/genai-demo-day Chalk and Elliot Marx won Best Technology at The GenAI Collective's Demo Day by showing how Chalk's feature store can be used to unify traditional machine learning features with generative AI for parsing unstructured data. A few weeks ago, Chalk won Best Technology at The GenAI Collective's Demo Night! In his demo, Elliot showed how Chalk achieves the best of both worlds by combining structured data with LLM analysis in a unified feature store. He used Chalk to ingest financial transactions from a traditional database and used Gemini to analyze unstructured memo lines to retrieve the merchant category, clean the memo, and find other meaningful information. Here's an abridged preview of how we implemented a feature pipeline for analyzing transaction data: import google.generativeai as genai import json from chalk import feature from chalk.features import features @features class Transaction: id: int # ... See full code for more features # Gemini features # Store completion as a separate feature so that we can iterate on # response parsing without retrying Gemini repeatedly completion: str = feature(max_staleness="infinity", default=default_completion) clean_memo: str category: str = "unknown" is_nsf: bool = False # NSF: insufficient funds is_ach: bool = False # ACH: direct deposit model = genai.GenerativeModel(model_name="models/gemini-1.5-flash-latest") @online async def get_transaction_classification(memo: Transaction.memo) -> Transaction.completion: return model.generate_content( textwrap.dedent( f"""\ Please return JSON for classifying a financial transaction using the following schema. {{"category": str, "is_nsf": bool, "clean_memo": str, "is_ach": bool}} All fields are required. Return EXACTLY one JSON object with NO other text. Memo: {memo}""" ), generation_config={"response_mime_type": "application/json"}, ).candidates[0].content.parts[0].text @online def get_structured_outputs(completion: Transaction.completion) -> Features[ Transaction.category, Transaction.is_nsf, Transaction.is_ach, Transaction.clean_memo, ]: """Given the completion, we parse it into a structured output.""" body = json.loads(completion) return Transaction( category=body["category"], is_nsf=body["is_nsf"], is_ach=body["is_ach"], clean_memo=body["clean_memo"], ) python Writing a feature pipeline with Chalk is as easy as writing Python. Chalk makes prompt engineering simple because Chalk handles passing structured data to your LLM prompts and caches responses to prompts indefinitely to reduce API costs. You can find Elliot's full demo on GitHub. It was a great night all around and we always love hanging with other members of the ML/AI community in San Francisco! # Building PartiQL Support at Chalk source: https://chalk.ai/blog/partiql A deep dive into how we built support for PartiQL and DynamoDB on Chalk by using DuckDB's SQL parser. PartiQL is a SQL dialect that can be translated into fast, native DynamoDB queries. We wanted to offer a way for our DynamoDB users to run PartiQL queries against their DynamoDB data sources, but there was one small wrinkle: DynamoDB's PartiQL dialect is too limited for use in Chalk SQL. In this blog post, we’ll discuss how we worked around these limitations to provide full support for customers using DynamoDB data sources with Chalk. We’ll dive into the details of how Chalk uses abstract syntax trees (ASTs) to parse and transform SQL. Background Chalk is a machine learning platform that lets you define your features (also known as signals or dimensions) in idiomatic Python. Feature extraction pipelines are automatically built from “resolvers”, which describe how to compute these features. Here's an example where we define features representing facts about our users and retrieve raw feature data from DynamoDB: from chalk.features import features @features class User: id: str name: str age: int python With Chalk, you can define SQL file resolvers, which populate features from your database. This SQL file resolver queries a DynamoDB instance to retrieve these values: -- type: online -- resolves: user -- source: our_dynamodb SELECT id, name, age FROM users; sql Chalk builds an internal graph of your features and resolvers, so it understands that queries for `User.id`, `User.name`, and `User.age` should invoke this SQL query. Importantly, Chalk implicitly maps the names of these columns to the names of features. If your database's column names are not an exact match for your Chalk feature names, we generally advise you to alias them using `AS`: -- type: online -- resolves: user -- source: our_dynamodb SELECT id, full_name AS name, age FROM users; And here's our wrinkle: PartiQL does not support `AS` in `SELECT` clauses! We considered several options: 1. Tell our users to rename their features to match their DynamoDB column names. - Problem: We cannot ask customers to rename columns and modify existing infrastructure just to use Chalk. 1. Add a way to define the name mapping within the file's metadata comments, such as with `-- rename: foo bar`. - Problem: Metadata comments are fine for Chalk-specific concepts, but they felt too magical in this situation. 1. Add a way to set an alias for the feature within Python code (e.g., `name: str = field(resolver_alias="full_name")`). - Problem: Putting the alias in the feature definition means that the SQL file resolver is not an independent source of truth about the feature mapping, which is confusing. 1. Find a way to allow `AS` within PartiQL resolvers - 🏆 This solution is intuitive and maintains consistency with other SQL data sources. This will require additional work under the hood, but is the best experience for our users. Although it was the most complex option, we decided the last option would be the most consistent and intuitive user experience. Our goal was to parse PartiQL queries and directly edit the abstract syntax trees (ASTs), and ultimately come up with a flavor of PartiQL that allows `AS` aliasing. Here's an overview of the steps we took: 1. Parse PartiQL queries to retrieve a mapping from column names to aliases 1. Run the query as valid PartiQL against DynamoDB 1. Apply the aliases to the query results Parse PartiQL queries First, we need to parse queries so that we can find the `AS alias` parts. An initial candidate for this type of work was SQLGlot, which is capable of creating the AST, modifying it, and transpiling it into any dialect we want. However, SQLGlot requires the bulk of work to occur in Python, where we would be subject to GIL-bound computation, which dramatically reduces performance. We instead chose to use DuckDB's SQL parser, which gave us access to an AST in C++. On the flipside, we have to do the work to transpile the AST back into a SQL string. We decided this tradeoff was worthwhile for improved performance. Here's how we run DuckDB’s parser on PartiQL queries to get an AST: duckdb::Parser p; try { p.ParseQuery(sql); } catch (const duckdb::ParserException &e) { ... } cpp ASTs are a data structure that represent the structure of code. They're a fundamental building block for writing compilers and executing code for programming languages. They make it easier for us to programmatically retrieve the properties we care about without writing impossible regular expressions. For example, for the following SQL query: SELECT id, full_name AS "name", FROM users WHERE status="active"; We would generate this AST: Starting from the root `SELECT` node, its child nodes include nodes for the `SELECT` clause, `FROM` clause, and `WHERE` clause. Within `SELECT_LIST`, there are nodes representing the `id` column as well as the `full_name` column. The `full_name` column node contains the `alias` property we care about. Here’s an abridged look into the `COLUMN` object in the DuckDB AST: namespace duckdb { class BaseExpression { public: string alias; virtual string ToString() const = 0; ... } class ParsedExpression : public BaseExpression {...} class ColumnRefExpression : public ParsedExpression { public: string column_name; ... } } We iterate over the `SELECT_LIST` node and retrieve the aliases. For aliased column references, we store the mapping of the column name to the alias. Finally, we remove the alias from the AST. In our example, we traverse the AST and store the mapping from the `full_name` column to its intended `alias`, `name`, and remove the `alias` value from `ColumnRefExpression`. Run the query Now that we have stored the column mapping and edited our AST to remove the aliases, we’re ready to convert our AST back into a SQL string. Unlike the previously-mentioned SQLGlot, the DuckDB AST was not created with the intention of converting back into SQL; it was created to query DuckDB. Because of this, just calling the built-in `ToString()` method, which converts queries into one specific dialect of SQL, does not always work. Different dialects have different semantics. Examples of differences include: - Different types of brackets for array literals: `(1,2,3)` vs `[1,2,3]` - Different quotes for column identifiers: `"foo"` vs ``foo`` - Different reserved keywords To get around this, we created a custom method that takes a SQL dialect and a DuckDB AST and returns the correct SQL string. Here’s a peek into what that looks like: namespace chalk_internal::sql { template std::string DialectWriter::query_select_node_to_string(const duckdb::SelectNode &expr) { std::stringstream result; // CTEs ... // SELECT result << "SELECT "; ... for (size_t i = 0; i < node.select_list.size(); i++) { if (i > 0) { result << ", "; } auto select_expr_as_str = expression_to_string(*node.select_list[i]); result << select_expr_as_str; if (!node.select_list[i]->alias.empty()) { result << " AS "; result << write_with_quotes(node.select_list[i]->alias, DialectInfo::identifier_quote); } } // FROM, WHERE, GROUP BY, HAVING, QUALIFY, etc. ... return result.str(); } } Using this, we're able to produce a valid PartiQL query: SELECT id, full_name FROM users WHERE status="active"; Add the aliases to the results Finally, we send our modified query to DynamoDB and parse the results. We rename the columns, and we’re done! outcome = client->ExecuteStatement(request); if (!outcome.IsSuccess()) { // Error handling } const auto &result = outcome.GetResult(); // Populate our record batch builder with the result std::shared_ptr result_builder = create_result_builder(result); auto result_record_batch = result_builder->Flush().ValueOrDie(); // Rename the columns! auto renamed_results = result_record_batch->RenameColumns(mapping_info).ValueOrDie(); return renamed_results; Conclusion Being able to query a variety of different data sources is critical to Chalk’s mission of delivering a bleeding-edge feature engineering experience. Supporting and extending PartiQL to query DynamoDB allows customers with large AWS ecosystems to access their data with very little overhead, without compromising on developer experience. If building low latency products with great developer experience sounds exciting to you, check out our open roles at chalk.ai/careers. # Quarterly Product Update: Summer source: https://chalk.ai/blog/product-update-july-2024 New product releases and changelog updates from Chalk for summer 2024. I'm excited to share the remarkable work our engineering team has shipped these past few months. I've compiled everything here in one place — future product updates will appear on this blog, so watch this space and follow our changelog to stay updated! What's new Chalk gRPC We shipped a gRPC engine for Chalk that improved performance by at least 2x through improved data serialization, efficient data transfer, and a migration to our C++ server. You can now use `ChalkGRPCClient` to run queries with the gRPC engine and fetch enriched metadata about your feature sets and resolvers through the `get_graph` method. Use SQL in offline queries With ChalkPy v2.38.8 or later, you can pass spine sql queries to offline queries instead of passing input data directly. Chalk will run your query on your offline data store and the resulting rows will be used as input to the offline query. Chalk will compute an efficient query plan to retrieve your SQL data without requiring you to load the data and transform it into input before sending it back to Chalk. Here's an example: output = chalk_client.offline_query( spine_sql_query=f""" SELECT t.txn_time AS ts, t.seller_id AS "seller.id", t.buyer_id as "buyer.id", t.amount as "txn.amount", t.payment_type as "txn.payment_type" FROM transactions AS t WHERE t.update_at >= {now - timedelta(days=30)} """, outputs=[ Seller.id, Buyer.id, Buyer.account_created_date, Txn.payment_type, # Computed in the seller namespace from the 'seller.id' spine feature. Seller.recent_transactions_volume, # Computed in the buyer namespace from the 'buyer.id' spine feature. Buyer.total_spent_last_30d, # Passed through from the SQL query Txn.amount, ], ) python Role-based access control for data sources and features We expanded the functionality of our service tokens to enable role-based access control (RBAC) at both the data source and feature level. On the datasource level, you can now restrict a token to only access data sources with matching tags to resolve features. On the feature level, you can restrict a token’s access to tagged features either by blocking the token from returning tagged features in any queries but allowing the feature values to be used in the computation of other features, or by blocking the token from accessing tagged features entirely. Track incremental resolver status via CLI You can now get the current progress state for your incremental resolvers by using the incremental status command: chalk incremental status --scheduled_query get_some_data__daily ✓ Fetched resolver progress state Resolver: N/A Query: run_this_query_daily Environment: chalk12345 Max Ingested Timestamp: 2024-07-01T16:01:46+00:00 Last Execution Timestamp: 2024-07-01T00:01:27.421873+00:00 Chalk deployment tags You can now add tags to your deployments. Tags must be unique to each of your environments. If you add an already existing tag to a new deployment, Chalk will remove the tag from your old deployment. Use the `--deployment-tag` flag: chalk apply --deployment-tag=latest --deployment-tag=v1.0.4 Heartbeat monitoring for long–running queries We now have heartbeating to poll the status of long-running queries and resolvers, which will now mark any hanging runs that are no longer detected as "failed" after a certain period of time. Miscellaneous improvements - We have integrations for Trino and Spanner as data sources. - Search and filter features in the Chalk feature catalog by their tags and owners. - Windowed resolvers have expanded to allow for hourly cadences. - SQL file resolvers now check your target return columns for typos, and suggest the closest named features. - Failed annotation parsing raises a type error with a more helpful error message. - SQL resolvers have improved error reporting for failures related to type conversion (e.g., if your resolver selects an int column, but the feature’s type is string). # Feature Store at Work source: https://chalk.ai/blog/fraud-risk-case-study Machine learning has long been deeply embedded in the field of fraud detection. A deep dive into how feature stores empower ML teams to build fraud models. Machine learning has long been deeply embedded in the field of fraud detection. Data science and engineering teams continue to develop increasingly sophisticated models for detecting fraud in real-time, which is vital for many industries, especially fintech. Production machine learning platforms power a majority of financial interactions across the world, making the infrastructure on which these models and platforms rely imperative. At the heart of machine learning infrastructure is the feature store. Feature stores enable data scientists, data engineers, and software engineers to extract accurate, real-time features, develop and iterate on machine learning models, and maintain historical stores of features past for training and compliance purposes. In this article, we will walk through a simplified version of a machine learning platform targeting fraud risk detection to expose how each component of the feature store fits together to facilitate a better fraud risk model. For a detailed dive into the architecture of a feature store, please refer to What is a Feature Store? High-Level Architecture The feature store lives between raw data sources, machine learning models, and applications. The schema of a feature store is called a registry, which defines the names and types of features. Once the features are defined in the registry, the core data flow determines how the features are computed. On top of that lives the monitoring and observability system that exports metrics used to detect and prevent problems like feature drift. The core data flow of a feature store can be divided into three parts. 1. ETL (Extract, Transform, Load): First, we load data from our raw data sources into the feature store. Our features can be directly loaded from the sources or require some computation based on the schema definitions in the registry, so we extract, optionally transform, and then load data into our stores. 1. Store: Within a feature store there are two forms of storage–-an online store and an offline store. The online store acts as a cache, serving recently or freshly computed real-time feature values with low latency. Meanwhile, the offline store handles large data volumes, usually storing historical feature values used for machine learning model training and data governance purposes. 1. Serve: The feature store serves data in response to queries. Depending on the query specifications, the data served may either be previously stored data living in the online or offline store, or it could be freshly computed on-demand feature values that reflect real-time data. The full flow of data in a machine learning platform is illustrated below, from the data sources, to the feature store, to the features served and used in the application. To illustrate how a feature store enables production machine learning platforms to handle fraud detection use cases, we will step through each portion of this data flow, starting with the raw data sources. Raw Data Sources A feature store integrates with your existing architecture. The first step in setting up a feature store for your machine learning platform is to define which raw data sources you want to connect. For this fraud detection model, we will use user and financial transaction data. First, we define a Kafka streaming source for live transaction data, as well as a Snowflake batch data source for enriched, historical transaction and user data. from chalk.sql import SnowflakeSource from chalk.streams import KafkaSource kafka = KafkaSource(name="txns_data_stream") snowflake = SnowflakeSource(name="user_db") python Then, we can use a real time data source, Risk API, to fetch live user risk scores. Once we have authenticated our feature store to each of these data sources, we’re ready to start defining features! Feature Store To construct our feature store, we can begin by defining the registry, which serves as a framework on top of which our core data flow will live, and finally define key alerting threshold for observability. Feature Store: Registry For our fraud use case, we will focus on two feature sets: User and Transaction. We define these feature sets, along with some resolver definitions for features, in the code below. from chalk.features import features, FeatureTime from chalk.streams import Windowed, windowed @features class User: id: int first_name: str last_name: str email: str address: str country_of_residence: str # these windowed features are aggregated over the last 7, 30, and 90 days # avg_txn_amount computes the average of transaction.amount over each time period # num_overdrafts computes the number of transactions where # transaction.is_overdraft == True over each time period. avg_txn_amount: Windowed[float] = windowed("7d", "30d", "90d") num_overdrafts: Windowed[int] = windowed("7d", "30d", "90d") risk_score: float # transactions consists all Transaction rows that are joined to User # by transaction.user_id transactions: "DataFrame[Transaction]" @features class Transaction: # these features are loaded directly from the kafka data source id: int user_id: User.id ts: FeatureTime vendor: str description: str amount: float country: string is_overdraft: bool # we compute this feature using transaction.country and # transaction.user.country_of_residence in_foreign_country: bool = _.country == _.user.country_of_residence Several features can be copied directly from the raw data source. We also use expression notation to define simple features, such as `transaction.in_foreign_country`. For the remainder of our features, we still need to define resolvers. We dictate how to load and transform data from our data sources into feature values using SQL resolvers or Python resolvers. The following SQL resolver loads existing columns from our Snowflake data source and maps them to the corresponding User features. -- resolves: User -- source: user_db select id, email, first_name, last_name, address, country_of_residence from user_db mysql The following Python resolver loads transactions from our Kafka streaming source and maps them to Transaction features. from pydantic import BaseModel from datasources import kafka from chalk.streams import stream # Pydantic models define the schema of the messages on the stream. class TransactionMessage(BaseModel): id: int user_id: int timestamp: datetime vendor: str description: str amount: float country: str is_overdraft: bool @stream(source=kafka) def stream_resolver(message: TransactionMessage) -> Features[ Transaction.id, Transaction.user_id, Transaction.timestamp, Transaction.vendor, Transaction.description, Transaction.amount, Transaction.country, Transaction.is_overdraft ]: return Transaction( id=message.id, user_id=message.user_id ts=message.timestamp, vendor=message.vendor, description=message.description, amount=message.amount, country=message.country, is_overdraft=message.is_overdraft ) We now have resolvers to compute all feature values that match data in our raw data sources. All that’s left is to define the resolvers for our remaining features. from chalk import online, DataFrame from kafka_resolver import TransactionMessage from risk import riskclient @online def get_avg_txn_amount(txns: DataFrame[TransactionMessage]) -> DataFrame[User.id, User.avg_txn_amount]: # we define a simple aggregation to calculate the average transaction amount # using SQL syntax (https://docs.chalk.ai/docs/aggregations#using-sql) # the time filter is pushed down based on the window definition of the feature return f""" select user_id as id, avg(amount) as avg_txn_amount from {txns} group by 1 """ @online def get_num_overdrafts(txns: DataFrame[TransactionMessage]) -> DataFrame[User.id, User.num_overdrafts]: # we define a simple aggregation to calculate the number of overdrafts # using SQL syntax (https://docs.chalk.ai/docs/aggregations#using-sql) # the time filter is pushed down based on the window definition of the feature return f""" select user_id as id, count(*) as num_overdrafts from {txns} where is_overdraft = 1 group by 1 """ # the cron notation allows us to run this resolver on a schedule # this resolver is currently set to run daily # https://docs.chalk.ai/docs/resolver-cron @online(cron="1d") def get_risk_score( first_name: User.first_name, last_name: User.last_name, email: User.email, address: User.address ) -> User.risk_score: # we call our internal Risk API to fetch a user's latest calculated risk score # based on their personal information riskclient = riskclient.RiskClient() return riskclient.get_risk_score(first_name, last_name, email, address) We have now completed defining our registry! Next, we can choose the storage component of our feature store. Feature Store: Core Data Flow The registry defines the schema for the first and last stage of the core data flow in a feature store–namely how we ETL (extract, transform, and load) data into the store and what features we will serve. Then the missing piece at the center is choosing which storage providers to use for the online and offline store. There are a few common choices optimized for performant cache reads for the online store and bulk storage for the offline store shown below. For this use case, let us use DynamoDB for the online store and Snowflake for the offline store. This completes the core data flow of the feature store! Feature Store: Monitoring and Observability Observability is required for all production systems, and feature stores are no exception. Feature stores provide a number of metrics that are useful for monitoring data quality. Given that we have a streaming source and a scheduled resolver, we can choose to monitor a few key metrics: - `resolver_high_water_mark`: we would want to determine the latest time stamp up to which we have loaded data from our streaming data source to ensure that we are not missing transaction data and to detect any issues with the stream. - `cron_feature_writes`: we can track the number of features written by our scheduler resolver to ensure that we’re consistently updating user risk scores with the latest data from our internal Risk API. - `feature_value`: we track the distribution of feature values over time to detect issues like feature drift, where small changes in accuracy of feature values can result in big impacts to our risk models. Thus, having defined our registry, core data flow, and observability metrics, we have all we need within our feature store to start to use it for our fraud risk detection model and application. Application: Querying the Feature Store Having configured the raw data sources and the different components of the feature store, we can now fetch feature values on demand through online and offline queries. Generally, online queries are used to fetch real-time feature values for live inference models and serving applications, whereas offline queries are used to query bulk and/or historical data, often for the development and training of machine learning models. For our fraud detection use case, the first step is querying features that we will use to train our machine learning models. Let’s say for a list of users, we want to query the risk scores, as well as the average transaction amounts and number of overdrafts in the last 7 days. from chalk.client import ChalkClient from src.models import User user_ids = [1093485493, 1093495732, 1029436728] dataset = ChalkClient().offline_query( input = { User.id: user_ids, User.ts: datetime.now() * len(user_ids), }, output = [ User.country_of_residence, User.risk_score, User.avg_txn_amount["7d"], User.num_overdrafts["7d"], ], recompute_features = True, ) To understand what happens when we run this query, we should think about how data is populated into the offline store. When we define resolvers, we can define them as `@online` or `@offline` based on how the features will be computed. Online resolvers are usually run on-demand, while offline resolvers are run to pre-compute features. A resolver will run either on a cron schedule, when it is triggered either manually, or by an orchestrator like Airflow. During a query, the feature store will use the resolver definitions to create a plan of which computations to run and in what order to determine all the outputs queried for the given inputs. Online queries may load data that is already cached in the online store, and offline queries may load already computed data from the offline store, but both types of queries can also serve real-time computed-on-demand feature values. In our query above, because we request `recompute_features=True`, the feature store recomputes all the feature values, so we have the most up-to-date data. So, the feature store will hit the Risk API and load and recompute feature values from our Snowflake (our batch data source) to determine the values for all the output features, which are served to the user to use in training and stored in the offline store. This data flow is illustrated below. The output for this query would look something like this: We can use these features to train our fraud risk model. To obtain the latest training data, we can schedule offline queries to ingest the latest relevant feature values into the offline store daily. Then, to use the fraud risk model to make real-time inference decisions to determine whether incoming transactions might be fraudulent, we can use an online query to fetch the relevant features for a user. from chalk.client import ChalkClient from src.models import User user_features = ChalkClient().query( input = {User.id: 1093485493}, output = [ User.id, User.first_name, User.last_name, User.country_of_residence, User.risk_score, User.avg_txn_amount["7d"], ] ) As a result of this query, we would get the following: { user.id: 1093485493 user.first_name: "Melanie" user.last_name: "Chen" user.country_of_residence: "USA" user.risk_score: 0.38 user.avg_txn_amount_7d: 31.23 } Since we want the online query to be low latency, we can also ingest the relevant feature values into the online store through a daily scheduled resolver run. Then, to determine whether an incoming transaction is fraudulent, we could run this online query to fetch the relevant user data, pass the user data and incoming transaction data into the model, and return a fraud risk decision in our application. The full flow of data for calculating the real-time fraud risk of a transaction looks like this. Thus, with our machine learning-powered application serving live fraud risk scores as computed using the features from our feature store, we have stepped through each component of a machine learning platform with a feature store. Conclusion In this case study, we demonstrate how a feature store can enable the development and deployment of machine learning models for real-time fraud risk detection in financial applications. Feature stores enable better collaboration between teams and the development of better machine learning models and products by centralizing the definition of key machine learning features and datasets. Fraud risk detection is just one of countless use cases where production machine learning teams can incorporate a feature store to improve their platform. Chalk offers a data source agnostic feature store, with clients in multiple programming languages, and lightning-fast deploys and queries. Developers love Chalk’s best-in-class feature store. To see for yourself, please book a demo. # Introducing Chalk NY source: https://chalk.ai/blog/introducing-chalk-ny Today we’re thrilled to announce Chalk's expansion with the opening of a New York City office. We’re thrilled to share that we’ve opened a brand new office in New York City. Our home base on the East Coast is in Flatiron, a few blocks from Madison Square Park. This expansion allows us to tap incredible East Coast talent and accelerate how Chalk helps developers ship production-grade machine learning. This move enables us to combine our world-class Bay Area team with the engineering, ops, and GTM talent in the Northeast. As Chalk continues its mission to empower ML and data teams, we’re excited to create a hub in New York and help foster the burgeoning AI and MLOps community here in the city. Chalk NY already has a small, but mighty, go-to-market org. With our expanded presence, we’re excited to grow our go-to-market team and build our engineering presence. If you are energized by the idea of defining Chalk’s culture across coasts and working alongside some of the world’s sharpest, we’d love to hear from you. Founding software and machine learning engineers, please apply! You can explore open roles at chalk.ai/careers. # CB Insights AI 100 source: https://chalk.ai/blog/cb-ai-top-100-2024 Today we're proud to announce CB Insights recognized Chalk alongside Databricks in its 2024 AI 100 as one of the leading ML / AI Development Platforms globally. I’m excited to share that today, CB Insights named Chalk to their eighth-annual list of the world’s 100 most promising AI companies. We’re thrilled to be recognized alongside Databricks as one of the leading Machine Learning and AI Development platforms globally. From the beginning, we set out to empower ML teams with a best-in-class developer experience which makes this recognition especially meaningful to us. With Chalk, developers have the necessary building blocks to ship production-grade machine learning and generative AI – a compute engine, an LLM toolchain, a feature store, integrated monitoring, and branches for experimentation. We’ve been excited to see industry leaders like Vital, Ramp, Melio and Whatnot use Chalk to power a wide range of complex predictions: detecting disease in real-time, fighting fraud, and transforming e-commerce. Across the board, companies leverage Chalk for critical use cases that require integrating the freshest possible data to make real-time decisions – a place where incumbent tooling is extremely limited. CB Insights shared with us that they selected winners based on team strength, investor strength, industry partnerships, patent activity, CB Insights’ data on deal activity, and proprietary Mosaic Scores. CB’s research team also conducted interviews with software buyers and reviewed hundreds of Analyst Briefings. We’re really honored to be included here. “AI is taking off at lightning speed, and it’s not just big tech companies at the forefront of it,” said Deepashri Varadharajan, director of AI research at CB Insights. “Our AI 100 winners…are pushing the boundaries of AI in everything from game development and battery design to agentic AI systems.” We’d love to hear about what you’re building next with data and see if there are some ways we might be of help. Please reach out to request a demo here. # $10M Seed Round to Accelerate ML source: https://chalk.ai/blog/announcing-chalk-10m-seed-funding We’re proud to announce we’ve raised $10 million in seed funding led by General Catalyst, Unusual Ventures, and Xfund to build the data platform for machine learning and generative AI. It’s an exciting day here at Chalk. We’re proud to announce we’ve raised $10 million in seed funding led by General Catalyst, Unusual Ventures, and Xfund to build the data platform for machine learning and generative AI. When my co-founder Elliot graduated from Stanford, his first job was at Affirm where he was tasked with building a data platform to support their buy-now-pay-later decisions. It was immediately clear that existing tools had challenging and cumbersome developer experiences – and didn’t support real-time use cases. He experienced that again trying to leverage incumbent solutions at his startup with Andy, and yet again at Credit Karma after they were acquired. Each time he wound up building his own solutions. Elliot sometimes jokes that building data infrastructure is the only thing he’s ever done (or knows how to do). Of course, that’s not quite true – but it is fair to say we’re intimately familiar with the painful workflows required to launch production-grade machine learning. Instead of relying on dated technologies like Spark, we built Chalk from the ground up to provide a best-in-class experience for our users – the experience we’ve always dreamed of. Chalk makes it simple to integrate new features, backtest changes, and deliver real-time machine learning and generative AI. Andy often says it best – Chalk's infrastructure is not just about powering decisions; it's about empowering developers. We deeply understand the pain points and workflows that make machine learning challenging, and Chalk was built to ease that pain. We’re very proud of the product we launched last year, and it’s been thrilling to see companies making some of the world’s most complex decisions leverage Chalk. Ramp, the ultimate platform for modern finance teams, selected Chalk to power core lending and fraud models – Chalk has become a critical component of our Risk Intelligence Platform. It expanded Ramp's capabilities with online machine learning and enabled us to scale safely by powering our transaction fraud model and credit underwriting process. Ryan Delgado, Staff Software Engineer, Data Platform, Ramp Other fintechs including Melio, Mission Lane, Pipe, and MoneyLion also leverage Chalk for critical use cases that require integrating the freshest possible data to make real-time decisions. Whatnot, the largest independent live shopping platform in the US, also uses Chalk as its machine learning feature platform to support recommendations, ranking, and fraud use cases for its buyers and sellers. In the healthcare space, Vital, a leading AI-driven digital health company, depends on Chalk to route ambulances based on hospital wait times and diagnosis sepsis in critical cases. Emergency room decisions can mean life or death for patients. Chalk powers these most critical predictions for us. We evaluated a wide range of technical solutions and Chalk stood out for having the very best developer experience at the most competitive cost. Te Riu Warren, CTO, Vital We’re incredibly fortunate to be supported by an amazing group of customers, investors, colleagues, and friends. We’ve been given the rare chance to build on lessons from our previous companies and are excited about the road ahead. Want to experience Chalk for yourself? Please book a demo. Also check out our interview with The Information! P.S. - We’re always looking to add incredibly talented people to our team. We’d love to hear from you – please check out our open roles. # Chalk Completes SOC 2 Type 1 Report source: https://chalk.ai/blog/soc-2 Chalk is proud to announce the availability of our SOC 2 Type 1 Report. At Chalk, we understand that our customers are trusting us to process their most important data assets. As part of our commitment to security and trust, we’re proud to announce that we have successfully completed our SOC 2 Type I audit. The SOC 2 audit is one of the highest recognized standards of information security compliance in the world. It was developed by the American Institute of CPAs (AICPA) to allow a third-party auditor to validate a service company’s internal controls with respect to information security. In order to prepare our SOC 2 Audit Report, Chalk worked with a third-party auditor to review our internal controls including but not limited to data security, network configuration, change management, access control, backup and disaster recovery and security incident response plans. SOC 2 is just one component of our security program. We are committed to continually improving our information security program and obtaining further certifications to support our customers' security and compliance needs. You can read more about our security program here, or reach out to us at security@chalk.ai for a copy of the report. # Get AI and ML with context into production. Fast. source: https://chalk.ai/capabilities/forward-deployed-engineering Forward-deployed Engineering at Chalk Book Architecture Review Forward-deployed Engineering Chalk’s forward-deployed engineers partner with your team to take critical AI workloads from architecture to production. Your AI roadmap shouldn't wait. Chalk was the only integrated compute and context engine for agents... What would have taken months took weeks. We continue to be impressed by the flexibility and scalability of the system, and by the willingness of the Chalk team to work with us to get even more value out of it. The Chalk team worked closely with us to optimize performance, solve edge cases, and even build features we needed. From serving context at the moment of decision to fine-tuning and securing agents, Chalk’s forward-deployed engineers bring the platform expertise and engineering capacity to help your team turn ambitious projects into production systems. Get technical guidance when you need it. Chalk’s forward-deployed engineers work with your AI, data, and business teams to define the workload, evaluate the tradeoffs, and map the path to production. Get help with hard decisions as you build, like what it takes to meet your latency, reliability, and security requirements. On-demand expertise Let Chalk take a workload to production. Bring us a high-priority AI/ML workload and we’ll help take it from design through deployment. We can fine-tune, evaluate, and deploy models using your data and infrastructure. A turnkey path from idea to custom AI deployment. Managed sprints Extend your team with Chalk engineers. Chalk engineers work alongside your engineers in your repo and cloud. They pair on implementation, open pull requests, and help build and operate production AI systems. Your team gets additional engineering capacity. Side-by-side capacity How we partner with your team Choose the engagement that fits your workload and your needs. Chalk's FDEs are the engineers who built Chalk working directly with you. Platform experience From Putnam finalists, PhDs, former Jane Street quants, to Rust compiler contributors. Technical experience They know how to connect live data to inference, design for latency requirements, and navigate the path to production. Production experience Why Chalk's forward-deployed engineers Chalk’s team brings specialized expertise to your toughest challenges, accelerating your path to production. AI workloads where real context matters. Chalk's forward-deployed engineers work with teams putting AI/ML at the core of the product - when you’re ready to tackle the bottlenecks between your data and your models at inference. Bring your highest-priority use case to a working session with a Chalk engineer. You’ll leave with: - A map of your current data and decision path - Where stale or disconnected data is limiting decisions - A scoped first workload with a clear plan to production - What it takes to run that workload in your cloud Make your AI/ML roadmap real. TALK TO US Book an Architecture Review Together, we’ll chart the path from your current architecture to a production system that makes decisions with live context. # Time traveling agent sandboxes source: https://chalk.ai/capabilities/compute Compute on Chalk submit_getstarted_form Request Access Request an account. Tell us what you’re building. Start now. Explore Docs Book Demo Compute Evaluate agents with knowledge cut-offs. Generate temporally-consistent context in your cloud. Compute that knows your data. Trusted by teams building the next generation of AI + ML Why Chalk Compute Chalk maintains a content-addressed image cache on every node in your cluster. Once an image lands on a host, every subsequent sandbox using it boots in under a second. Common base images are pre-warmed across the fleet, so cold starts are an exception, not a default. Scale-to-zero costs you nothing on the way back up. Fast. Sub-second cold starts on every workload. Every workload runs inside a gVisor-hardened sandbox — a user-space kernel intercepts syscalls before they reach the host. Each sandbox launches with its own OIDC-compliant cloud identity, scoped to that workload alone. Outbound traffic is restricted to a hostname or CIDR allowlist you control. A compromised agent stays inside its blast radius. Safe. gVisor isolation, OIDC identity, policy-bound egress. Sandboxes execute in your VPC, on nodes you provisioned, under IAM roles you control. Customer data never crosses to a third-party plane. Logs stay in your account. KMS keys, network policies, and audit trails are the ones your security team already operates. Chalk's metadata plane orchestrates the work; your data plane owns the data. Yours. Runs inside your AWS, GCP, or Azure account. Always. Designed to help AI teams deploy safely and securely. Production-grade by default. One platform. Every Agent. Built for production. Agents you trust. Fraud triage Loan pre-screening Underwriting Document analysis Customer support Financial analysis Data analysis Coding Sentiment analysis Deep research Sourcing One source of truth for every feature, online and offline. Feature Store Train, version, and roll out models on the same engine that serves them. Model Platform Sub-millisecond inference for production ML. Real-Time Serving Prompts, embeddings, evals, and tool-use as first-class features. LLM Toolchain Compute is one piece of the unified platform. Chalk Compute and the Chalk Context Engine deploy entirely inside your AWS, GCP, or Azure account. It all runs in your cloud. A high-level managed container that bundles image build, file upload, and lifecycle in a single Python class. Configure CPU, memory, GPU, secrets, volumes, and lifetime — then .run(), .exec(), and .stop(). Containers gVisor-isolated execution environments with their own filesystem, network namespace, and resource limits. CPU and GPU workloads, kernel-level multi-tenancy. Sandboxes Long-running, autoscaling HTTP services deployed from an image. Configure min and max replicas, target CPU utilization, and a port — Chalk handles the load-balanced endpoint and graceful scale-down. Scaling Groups Core Infrastructure A fluent Python builder for defining software environments. Content-addressed caching means rebuilds are near-instant and identical across every sandbox. Images Persistent, versioned file storage with copy-on-write semantics. Fork a volume to fan out parallel workloads against the same snapshot — no copies, no drift. Volumes FILESYSTEM Deploy any Python callable as a remotely invocable endpoint. Automatic scaling, no servers to provision, no Dockerfiles to maintain. Functions Durable, ordered task execution with retries and backpressure. The right substrate for agent loops, evals, and bulk inference jobs that have to actually finish. Function Queue FUNCTIONS THE PRIMITIVES Simple primitives. Endless possibilities. Chalk Compute exposes the small set of primitives you actually need to run agents and inference in production — scaling groups, sandboxes, containers, images, volumes, functions, and function queues — with the boring parts (caching, isolation, autoscaling) already solved. OPENAI AGENT SDK ANTHROPIC SDK LANGCHAIN LLAMAINDEX MCP SERVERS We don't care which harness you build on — Chalk Compute is the infrastructure underneath. Bring your own agent harness. Every workload runs under gVisor — a user-space kernel that mediates syscalls before they reach the host. Inside the sandbox, CapEff and CapBnd are zero, no-new-privs is on, and securebits (secure-noroot, secure-no-suid-fixup, secure-keep-caps) are locked. Root inside the container has nothing to escalate to. gVisor isolation The interfaces with the largest historical bug count are blocked entirely: io_uring, bpf, perf_event_open, userfaultfd, fanotify, and kexec_load all return permission denied. /dev/kcore, /dev/mem, and /dev/port don't exist; /sys/kernel is empty; host block-device mounts aren't visible. Kernel surface unreachable Every sandbox launches with its own OIDC-compliant cloud identity, scoped to that workload alone. When self-hosted on Kubernetes, no service account token is mounted into the workload — it cannot authenticate to the cluster API by default. Workload Identity Federation Restrict outbound egress to a hostname or CIDR allowlist; off-list traffic is dropped silently at the network layer. Raw packet sockets (AF_PACKET) and link-level admin (ip link) are blocked outright — a compromised agent can't sniff the wire or reconfigure interfaces. Network Policy Sandboxes call MCP servers through the gateway, authenticated by their workload identity. The gateway holds the real credential and proxies the call — agents get tool access without ever seeing the upstream key. MCP Gateway Per-session WireGuard tunnels with dynamically negotiated keys, scoped to a single workload. Connect sandboxes to each other or bridge to on-prem databases without exposing them to the public internet. WireGuard Tunnels SECURITY Hardened at the kernel. Locked at the network. Identified at the workload. Every sandbox runs under gVisor — a user-space kernel that intercepts syscalls before they reach the host. Inside the sandbox, root has no effective capabilities, no-new-privs is on, and the kernel interfaces with the worst historical bug count are sealed off. We run a probe suite against every build to verify the sandbox holds. Agents for Compliance / Fraud Triage / Loan Pre-Screening / Underwriting / Document Analysis / Customer Support / Financial Analysis / Data Analysis / Coding / Sentiment Analysis / Deep Research / Sourcing In a sandbox. In your cloud. READ THE DOCS TALK TO AN ENGINEER See what your team can ship when sandboxes, models, and agents all run on the same engine — inside your cloud. The latest at Chalk Explore more of Chalk’s data platform # Real-time feature serving for production ML source: https://chalk.ai/capabilities/real-time-serving Chalk delivers real-time feature serving with millisecond latency. Serve fresh, consistent features directly at inference for production ML. TALK TO AN ENGINEER REAL-TIME SERVING Serve production features at inference time with single-digit millisecond latency. Chalk’s query execution engine runs Python functions, queries databases, and calls APIs in real-time, enabling decisions on the freshest data without brittle ETL or stale caches. Trusted by teams building the next generation of AI + ML from chalk import online @online def get_username(email: User.email) -> User.username: username = email.split("@")[0] if "gmail.com" in email: username = username.split("+")[0].replace(".", "") return username.lower() python Why teams choose Chalk for real-time serving Built for real-time Chalk is a query execution engine that can run Python functions, query databases, and call APIs at inference time. Compute the freshest data without complex and stale values from reverse ETL. Python Resolvers Low latency by design Chalk parallelizes execution and builds an optimized query plan that avoids redundant computation, enabling sub-5ms end-to-end feature retrieval. QUERY PLANNER Deploy securely in your VPC Deploy Chalk directly inside your AWS, GCP, or Azure cloud account. Keep all data, features, and models within your VPC while operating under your own IAM, networking, and encryption standards. DEPLOY IN YOUR VPC Resolve feature values in single-digit milliseconds under peak load. Performance at scale Unify feature definitions across batch, real time, and inference without drift or duplicated pipelines. Consistent with training Use familiar tools like Pydantic, Numpy, and Pandas. Chalk runs your Python code at scale without DSLs or wrappers. Native Python support Define dependencies with Python signatures. Chalk orchestrates resolvers into efficient query plans across online/offline environments. Declarative pipelines Experiment faster. Chalk spins up deployments per branch for safe iteration and review. Iterate faster Run Python at native speed. Chalk uses Rust to parallelize fetches, push down ops, and multithread computations. Rust-powered runtime Chalk has transformed our ML development workflow. We can now build and iterate on ML features faster than ever, with a dramatically better developer experience. Chalk also powers real-time feature transformations for our LLM tools and models — critical for meeting the ultra-high freshness standards we require. \# get_username converted into a static expression lower( if_then_else( !=( strpos(str(email), str(gmail.com)), int(0) ), replace( element_at( split( element_at(split(str(email), str(@)), int(1)), str(+) ), int(1) ), str(.), str() ), element_at(split(str(email), str(@)), int(1)) ) ) One language and one system for training and inference. Your Python becomes the source of truth and is automatically optimized for real-time execution. Chalk: - Builds highly efficient query plans so you only compute and fetch what you need. - Parses and transpiles your Python into static native expressions that run at Rust speed. - Centralizes your ML logic in code, removing duplication and eliminating feature rewrites. LEARN MORE Accelerated execution with Python-to-Rust transpilation The latest at Chalk Talk to an engineer Compute fresh features. Serve them in single-digit milli­seconds. Unify your ML workflow. Explore more of Chalk’s data platform # Train models with production features source: https://chalk.ai/capabilities/training-data Generate consistent training datasets for ML. Chalk backfills data with point-in-time correctness, ensuring reproducibility and scalability. TALK TO AN ENGINEER TRAINING DATA Generate point-in-time correct training datasets from the same feature definitions you use in production and serve production features from your training data. Keep training and inference aligned and eliminate skew. Trusted by the teams building next generation of AI + ML Evaluate training with the same feature definitions and resolvers used online to eliminate skew and rewrites. Training-serving consistency Generate billions of rows of training datasets through distributed offline queries without provisioning infrastructure. High-throughput, large scale Evaluate features using only the data available at each historical moment to avoid leakage and ensure accurate training inputs. Point-in-time correctness Create feature branches, generate comparison datasets, and seamlessly promote winners to production. Branch-based experimentation Produce versioned datasets for reproducible results. Deterministic & reproducible Chalk is a key part of our underwriting pipeline, letting us test new code against production pipelines without disruption. We can quickly iterate on new features and enhancements, helping us offer more flexible capital products than anyone else. Point-in-time evaluation ensures features are computed using only the values available at each historical moment. This removes leakage, produces reproducible datasets, and allows offline experiments to reflect production behavior. Explore offline queries Generate point-in-time datasets from production features Training datasets generated with Chalk integrate directly into your model training workflow. Chalk preserves the schemas and evaluation semantics used in production so models train on the same feature logic they use at inference. Model training Integrate with your model training workflow from chalk import Now, online @online def get_age(birthdate: User.birthdate, now: Now) -> User.age: return (now.date() - birthdate).days // 365 Evaluate features as if it were any point in time. Train on historically accurate windows, then serve the same features in production with no drift or leakage. Use chalk.now Backfilling with chalk.now The latest at Chalk Talk to an engineer Train your models with production features Talk to an engineer and see how Chalk can power your production AI and ML systems. Explore more of Chalk’s data platform # Every feature has a single source of truth source: https://chalk.ai/capabilities/feature-store Chalk’s real-time feature store lets you define features once and compute them on demand across training, batch scoring, and real-time inference. TALK TO AN ENGINEER FEATURE STORE Chalk’s feature store powers production ML. Define features once, and Chalk computes them on demand for training, batch scoring, and inference. Features stay fresh and consistent across environments. Trusted by teams building the next generation of AI + ML Define once, use across training, batch, and serving. One feature definition everywhere Instead of just storing values, compute features at query time. Execution first Recompute features as they would have been at inference time. No leakage, no skew. Point-in-time correctness Experiment quickly and safely on branches. Iterate fast with branches Serve precomputed or on-demand features in sub-5ms at scale. Low-latency serving Built-in catalog, versioning, and metadata for auditability. Discoverability & governance Chalk has become a powerful addition to our ML infrastructure at Mission Lane, unifying and streamlining feature calculations for offline, batch evaluation, and live decisioning. @features class Transaction: id: int amount: float user_id: "User.id" @features class User: id: int fraud_score: float = feature(version=2) txns: DataFrame[Transaction] count_small_txns: int = _.txns[_.amount < 20].count() raw_credit_report: str = feature(max_staleness="30d") python Chalk isn’t just storage. It’s a query execution engine for your features. At request time, Chalk computes only what’s needed, from the freshest data available. EXPLORE ONLINE QUERIES One feature catalog The latest at Chalk Eliminate training-serving drift Talk to an engineer and see how Chalk can power your production AI and ML systems. Explore more of Chalk‘s data platform # Temporal context at scale source: https://chalk.ai/capabilities/temporal-aggregations Generate consistent training datasets for ML. Chalk backfills data with point-in-time correctness, ensuring reproducibility and scalability. TALK TO AN ENGINEER TEMPORAL AGGREGATIONS Define aggregations once and reuse them across batch, online, and real time. Trusted by teams building next generation of AI + ML Continuously update aggregations from streams or compute at the leading edge of your transactional sources for maximum freshness. Alway-fresh features Define aggregations once and reuse across batch, online, and real time without drift. Consistent across training and serving Incremental updates keep queries fast and memory predictable as data volume grows. Efficient at scale Aggregate across streams, warehouses, and databases through our feature engine. Join across streams Chalk’s performance directly affects the quality of our search and discovery models, which power everything from price flexibility to apartment ranking. The ability to call real-time features without dealing with stream complexity has been huge for us. text How temporal aggregations work Materialized aggregations Precompute incremental buckets that update as events arrive for fast, production-ready features. Cache and materialize features Windowed aggregations Run exact queries over any time horizon to rescan historical data. Define features over time ranges Streaming + materialized aggregations Apply materialization directly to live event streams. Features update continuously without needing to store raw data. Aggregate functions on streams The latest at Chalk Talk to an engineer Temporal context for your ML features Talk to an engineer and see how Chalk can power your production AI and ML systems. Explore more of Chalk’s data platform # Real-time context for LLM inference source: https://chalk.ai/capabilities/llm-toolchain Chalk’s LLM toolchain unifies prompt engineering, embeddings, vector search, and real-time inference so you can ship LLMs faster, all the way to production. TALK TO AN ENGINEER LLM TOOLCHAIN A unified interface to serve, evaluate, and optimize LLMs using structured and unstructured data with sub-5ms latency. Trusted by teams building the next generation of AI + ML \# Use structured output to easily incorporate # unstructured data in your ML pipelines class AnalyzedReceiptStruct(BaseModel): expense_category: ExpenseCategoryEnum business_expense: bool loyalty_program: str return_policy: int @features class Transaction: # Named prompts enable editing prompts directly # in the Chalk UI without redeploying code llm: P.PromptResponse = P.run_prompt( "analyze_receipt_with_prompt_from_chalk_dashboard" ) # Let Chalk manage injecting the # right features at inference time user_prompt: str = F.jinja( """ Analyze the following receipt: Line items: {{Transaction.receipt.line_items}} Merchant: {{Transaction.merchant.name}} {{Transaction.merchant.description}} """ ) python LLM Toolchain benefits Ship LLMs faster Develop, evaluate, and deploy prompts and models in one system, with minimal glue code. Built-in scheduling, streaming + caching Inject live structured features directly into your prompts, without ETL or batch jobs. Standardize LLM development Use versioned, parameterized prompts and completions as first-class objects in your stack. Scale without fragmentation Unify feature engineering, vector search, LLM inference, and monitoring on a single platform. LLM Toolchain Docs Experiment with prompts on historical data using branches. Chalk tracks outputs, computes metrics, and promotes winning prompts with one command. Prompt Engineering Deploy inference pipelines with autoscaling and GPU support. Write pre/post-processing in Python. Chalk handles the rest, including data logging and versioning. Model Inference Log and compare model outputs with quality metrics to pick the best prompt, embedding, or model—all versioned automatically in Chalk. Evaluations Use any embedding model with one line of code. Chalk handles batching, caching, and lets you safely test new models on all your data. Embedding Functions Run nearest-neighbor search directly in your feature pipeline. Use any feature as the query, and generate new features from search results. Vector Search Process and embed large files, docs, images, and videos at scale. Chalk handles batching, autoscaling, and execution with a fast Rust backend. Large File Support One Platform. One Toolchain. All the way to production. Chalk powers our LLM pipeline by turning complex inputs like HTML, URLs, and screenshots into structured, auditable features. We can serve lightweight heuristics up front and rich LLM reasoning deeper in the stack, catching threats others miss without compromising speed or precision. @features class ProductRec: user_id: Primary[User.id] user: User user_vector: Vector = embed( input=F.array_join( F.array_agg( _.user.products[ _.name, _.type == "liked" ] ), delimiter=" || ", ), provider="vertexai", model="text-embedding-005", ) similar_users: DataFrame[User] = has_many( lambda: ProductRec.user_vector.is_near( User.liked_products_vector ) ) - Retrieve structured features dynamically at inference time - Use Python (not DSLs) to define feature logic - Fetch real-time context windows with point-in-time correctness - Mix embeddings and features for fully grounded RAG workflows CHALK FOR AI ENGINEERS Connect your LLMs to the freshest data without ETL pipelines chalk_client.prompt_evaluation( evaluators=["exact_match"], reference_output="review.sentiment", prompts=[ "analyze_sentiment-v1", "analyze_sentiment-v2", P.completion( model="gpt-5.1-2025-11-13", messages=[ P.user_message( """Analyze the sentiment of this product review: Review text: {{review.text}} Rating from this user: {{review.rating}} Average rating for product: {{review.product.average_rating}} Average rating this user gives: {{review.user.average_rating}}""" ) ], ), ] ).to_pandas() - Write, version, and reuse prompts with structured parameters - Evaluate prompts and models using historical production data - Compare model performance on accuracy, latency, and token usage - Debug failures with end-to-end traceability and lineage - Deploy prompt + model bundles as artifacts with full observability NAMED PROMPTS Design prompts like you design software The latest at Chalk See how Chalk compiles prompt logic, feature queries, and completions into optimized inference pipelines. Explore more of Chalk‘s data platform # Build or Buy? Why Melio Picked Chalk source: https://chalk.ai/customers/melio Why Melio replaced their in-house feature store with Chalk to power risk decisioning on every B2B payment and opened feature development to 20+ contributors. Melio uses Chalk to power risk decisioning on every B2B payment they process, replacing a homegrown feature store that bottlenecked engineering with a managed, self-serve platform that scaled feature development from a handful of engineers to 20+ across their risk org. Melio, recently acquired by Xero, processes billions of dollars in B2B payments for over 100,000 small and medium-sized businesses in the United States. The company has raised $654 million, landed on the Forbes FinTech 50 and Cloud 100, and built an international team with over 200 engineers. Every payment that Melio processes carries risk such as fraud, compliance violations, and non-sufficient funds. If the risk system is slow, this impacts payment processing speed. If the risk system is wrong, Melio could lose money. For years, the team that owned risk infrastructure also owned a homegrown Python feature store. It worked on a small scale — until it didn't. The decision became simple: keep pouring engineering into infrastructure, or hand that problem to Chalk and put those hours toward the models where mistakes carry real consequences. They made the switch. Overview As Melio's payment volume grew further, its needs quickly outpaced what a basic feature store could support. Scaling-related upkeep absorbed most of the engineering team’s bandwidth. The bottleneck was felt most by the people closest to the models. Data scientists on the risk team were locked out. Only a handful of engineers who knew the system's internals could ship anything. On top of this, the risk team's requirements were demanding. Melio's compliance models require a high degree of accuracy to meet strict regulatory SLOs. Fraud models operate on probability and require a different kind of tolerance. Both had to coexist on the same infrastructure. Mush Kabalo, who runs Melio's Risk Platform team, had seen this movie before: "I worked at two companies that tried building a feature store internally. Getting results was always an uphill battle. When I came to Melio and worked with Chalk, it felt really easy." The Challenge Chalk replaced Melio’s in-house feature store entirely. Scaling, versioning, metadata management, caching, the offline store: all of it became possible when Melio moved to Chalk. The engineers who maintained the old system could finally focus on what actually attracted them to join Melio: building fraud and compliance models. Chalk absorbed the work that had been eating the team's time and gave Melio control over how risk assessments hit their databases. The team opted to run Chalk against DB replicas on a cluster sized for both application and Chalk traffic — protecting production while keeping state accurate. Chalk connected to all of Melio's data sources: MySQL, Postgres, Snowflake, and native enrichment APIs. This isn't a nice-to-have integration. Chalk powers risk assessment on every payment that moves through Melio. Fraud detection, compliance checks, non-sufficient funds evaluation: all on the same platform. Before Chalk, feature development lived with a small group of engineers. After Chalk, roughly 20 developers across the risk organization write and deploy features, including data scientists, without needing engineering permissions. That last part matters. Melio's data scientists know what features the models need. With Chalk, they can more easily shape new and existing models. Teams now add enrichments, modify features, and deploy changes without routing through the infrastructure team. Then something unplanned happened. Melio's data analysts started using Chalk to enrich product analytics events. Now application teams outside risk are asking the same question: should we query the database directly and write pipeline code around it, or just write features on Chalk? That conversation didn't exist a year ago. The Solution Chalk sits between the risk decision models and Melio's data. When a payment is initiated, the risk platform queries Chalk for features, passes them to the model, and returns a decision. - Inbound data sources: databases, data warehouses, and external APIs - Primary flow: Application triggers risk assessment → Risk platform queries Chalk → Chalk computes and serves features → Models return decision - Secondary flow: Kinesis events enriched through Chalk for product analytics → Output to Snowflake, Tableau, FullStory, Amplitude Chalk owns the feature layer. The risk platform orchestrates, querying Chalk for structured feature inputs and feeding them to models and, increasingly, to risk agents that assist with manual review and automated rule-based decisioning. Architecture Replacing Melio’s feature store changed how the team works day to day. Before Chalk, Melio's in-house feature store needed dedicated engineering maintenance and limited contributions from the team. With Chalk, the feature store is fully managed with zero internal headcount. 20+ developers and data scientists across the risk org now ship features independently, risk workloads run separately from production, and use cases have grown beyond risk into product analytics — with more application teams asking to get on. Outcomes Three years ago, Chalk replaced a broken feature store. Today it supports the system that decides whether payments clear, the layer that feeds product analytics, and the platform that non-technical teams are starting to build on. "The stability of Chalk has improved vastly for us. And given that our scale also grew, that means double the improvement." Melio’s next priorities are practical. More teams want in. Data quality monitoring matters now that 20+ people are writing features instead of four. And the risk agents that already consume Chalk features for manual review and automated decisioning are only going to get more sophisticated. Looking Ahead In-house feature store couldn't scale with payment volume Risk reliability outgrew what the homegrown system could guarantee Feature development was limited to a handful of engineers AWS Fintech Managed feature platform replaced the internal feature store team Mission-critical uptime for fraud, compliance, and NSF on every payment Self-serve feature development for 20+ engineers and data scientists Risk decisioning, Compliance # How Turo Built a Self-Serve ML Feature Platform for Search and Pricing with Chalk source: https://chalk.ai/customers/turo How Turo built a self-serve ML feature platform for search and ranking using Chalk’s compute-first feature store, reducing data overhead and enabling predictable, low-latency production performance. Turo standardized feature delivery across search, pricing, and risk using Chalk’s compute-first feature store, enabling faster iteration, predictable production performance, and a scalable foundation for real-time ML. Turo is the world’s largest car-sharing marketplace, with more than 300,000 host-owned vehicles across the US, Canada, the UK, France, and Australia. At this scale, machine learning systems are foundational infrastructure. They determine how guests discover cars, how owners price trips, and how trust is enforced across the marketplace. In a two-sided marketplace, small changes in relevance, pricing, or trust compound quickly. Those gains show up as higher liquidity, stronger conversion, and increased revenue. Turo applies machine learning across three critical areas: ranking vehicles in search results, providing dynamic pricing recommendations for hosts, and assessing risk. These systems span both batch and latency-sensitive workloads and operate at different points in the funnel, from high-traffic search requests to downstream risk decisions. Supporting this range of use cases requires a compute-first feature store that handles both batch and real-time feature computation, with low-latency serving in production. Slow iteration or inconsistent features don’t just hurt models. They slow feedback loops between supply and demand. Operating at global scale with hundreds of thousands of active vehicles, Turo sees outsized impact from even small improvements in search relevance or pricing accuracy. Search, pricing, and risk all have very different requirements, but they all depend on features being predictable and available in production. Overview Before Chalk, Turo did not have a unified feature store for production feature delivery. Feature data lived in multiple places. Machine learning engineers often had to comb through production databases to find viable data sources, then craft Airflow jobs and bulky pipelines to make feature data available for online use cases. Feature definitions and model logic were scattered across repositories and services, making it difficult to treat features as shared, reusable production assets. As a result, feature delivery became a constraint on how quickly Turo could adapt pricing recommendations, improve search relevance, or refine risk decisions as marketplace dynamics changed. Each new feature followed a different path to production. Shipping changes meant locating the right data, extending bespoke pipelines, and coordinating with multiple teams to provision infrastructure. The bottleneck wasn’t modeling. It was operationalizing features fast enough to keep pace with the marketplace. For a marketplace business, this friction directly affected how quickly Turo could learn from user behavior and improve liquidity, conversion, and trust across the platform. We had all the pieces of an ML platform, but no single, consistent way to productionize features. The Challenge Turo adopted Chalk to create a single, repeatable path for production feature delivery, using Chalk as its centralized feature store, owned by the ML engineering team and deployed in Turo’s own cloud. With Chalk, feature definitions live in one place and follow consistent patterns across different models. Features are defined around shared entities and computed through resolvers that encode how data is fetched, joined, or refreshed in production. ML engineers, MLOps, and data scientists can all directly define and iterate on features in Chalk, using the same abstractions and workflows. Chalk absorbed much of the feature-related data engineering work that previously required custom pipelines or cross-team coordination. The team developed standard workflows that cover roughly 80 percent of feature delivery needs, allowing feature development to happen collaboratively while the ML engineering team owns production ingestion, serving, and APIs. This allowed the ML engineering team to own feature delivery end-to-end, without relying on data engineering teams whose priorities were focused primarily on analytics rather than ML. At a platform level, this enables Turo to standardize on a small set of repeatable primitives: - Centralized, entity-based feature definitions shared across models - Resolvers that define how features are computed, ingested, or refreshed - An online feature store optimized for low-latency production access - Self-serve workflows owned by the ML engineering team @features(max_staleness="10d") class InternalAutomaticPricingMinMax: id: Primary[str] min_max: list[MinMax] def create_ap_min_max_resolver(config: SQLResolverConfig) -> None: sql = load_sql_template("automatic_pricing_internal_min_max.sql") for version, table_name in TABLE_MAP: if config.environment == "staging": cron = "0 20 * * 0" else: cron = None make_sql_file_resolver( name=f"get_{config.environment}_{version}_ap_internal_minmax", kind="offline", environment=config.environment, sql=sql.format( default_db_name=config.default_db_name, version=version, table_name=table_name, ), source=redshift, resolves=InternalAutomaticPricingMinMax, tags=f"ap_internal_minmax_{version}", cron=cron, ) python The example above shows a resolver loading data from a SQL table on a daily schedule. Before Chalk, moving data from analytics or data science into Turo’s production microservices required custom pipelines and coordination across teams. With Chalk, teams can populate a SQL table and trigger offline feature computation, making the data available to production services through a standardized feature interface, typically within a day. This shift reduced dependency on platform and data teams for most feature work. It also made production practices more consistent as the number of models and features increased. Chalk’s compute-first architecture allowed the team to define feature logic once and rely on Chalk to handle execution and serving, rather than maintaining bespoke pipelines per model. Chalk let us turn feature delivery into a self-serve workflow. We stopped waiting on other teams for every new feature and started shipping on our own cadence. The Solution Turo runs Chalk in its AWS cloud environment. The ML engineering team manages the platform day to day and works with the platform team for deeper infrastructure support when needed. Data ingestion patterns Most production features follow a SQL-driven workflow: - Data scientists or MLOps workflows generate tables using Airflow - Airflow triggers Chalk offline feature computation via SQL and Python resolvers - Chalk computes and loads features into the online store - Services request features at inference time through internal APIs This design positions Chalk as a feature engine focused on predictable, low-latency serving in production. Models and latency Chalk supports several core model families at Turo: - Pricing models, which are largely batch-oriented and have the highest feature volume - Vehicle search models, which are latency-sensitive and operate high in the funnel - Risk models, which tend to be lighter-weight and run later in the flow Based on their integration, Chalk provides consistent response times that allow engineers to design systems around expected feature retrieval latency. For our high-funnel search workloads, predictability matters. Architecture Standardizing feature delivery on Chalk improved both iteration speed and operational clarity. evenSplit - Time to ship a new production feature - Dependency on other teams - Feature delivery consistency - Production predictability Metric - At least two to three weeks when coordination was required - Frequent coordination with platform and data teams - Different workflows per model - Inconsistent due to fragmented pipelines Before Chalk - One week - Self-serve within the ML engineering team - Shared patterns cover most use cases - Predictable latency that teams can plan around After Chalk As model complexity increased, Chalk scaled with it. The platform absorbed more load as teams iterated faster, rather than becoming a bottleneck. As our models get more complex, Chalk gets used more, not less. Outcomes Turo is extending its feature platform toward streaming, using Chalk as the foundation for both batch and real-time feature delivery. As new models are introduced, feature work now defaults to Chalk first, with streaming ingestion becoming a natural extension of the same feature definitions and serving patterns already in production. This approach allows the ML team to pursue more advanced real-time use cases without reintroducing data engineering bottlenecks or disrupting existing production workflows. Chalk has shifted feature engineering at Turo from a coordination-heavy bottleneck into a scalable platform capability. The result is faster iteration today and a durable foundation for streaming and real-time ML tomorrow. What's Next Fragmented feature pipelines across production systems Slow time to production for new features Inconsistent feature workflows across models AWS Marketplace A single, centralized feature store replacing fragmented pipelines Self-serve feature delivery accelerating time to production Standardized, resolver-driven feature workflows shared across models Search and ranking, Dynamic pricing # Medely staffs critical healthcare in real-time with Chalk source: https://chalk.ai/customers/medely Medely uses Chalk's self-serve feature platform to seamlessly deploy ML models, enabling real-time pricing and matching for healthcare staffing. Medely uses Chalk to staff critical healthcare roles in real time, replacing a fragile batch pipeline with self-serve, real-time feature infrastructure that accelerates ML velocity across pricing and matching. Medely is the world's largest healthcare talent marketplace, connecting providers to a flexible workforce of 300,000 nurses and allied professionals. Since COVID-19, demand for healthcare gig work has surged: a 2022 Oliver Wyman survey found a 1400% spike in nurses moving to gig models. Medely is creating a trusted network where healthcare workers can find flexible opportunities and facilities can fill urgent shifts. As Medely matured, the company recognized that machine learning could optimize two key levers: - Dynamic pricing: Recommending charge rates that balance facility budgets with competitive compensation - Professional matching: Placing qualified professionals into the right jobs at the right time to maximize fill rate But without a proper experimentation-to-production pipeline, these optimizations remained out of reach. Medely chose Chalk for self-serve feature infrastructure that could scale with their ambitions. Within months, they deployed a charge rate recommendations model that drove a sharp increase in revenue, proving the value of ML and clearing the path to growth. Overview Healthcare staffing operates on compressed timelines. Jobs can be posted with as little lead time as 48 hours. Pricing and matching models need to react to marketplace conditions in real time. Medely assembled a batch pipeline from available tools: - Prefect - Redis - Inference Service Component - Orchestrated a daily batch job that queried Postgres/Snowflake - Stored pre-computed features (feature store) - Read features from Redis to make predictions Function This architecture created two main problems: 1. Models operated on stale data Since the job was daily, feature freshness was locked at 24 hours: too slow to react to rapidly changing facility demand and professional availability. 2. Every iteration required heavy manual work - Write a new feature - Modify feature logic - Add a data source - Modify schema ML task - Write separate Snowflake queries (including complex joins for relational features) - Redeploy Prefect pipeline across entire stack - Coordinate with data engineering + update pipelines - Provision infrastructure Infrastructure + development overhead It was quite clear to me that the ownership and the infra was not going to scale with the ambitions of the number of models we wanted to deploy. To scale their ML use cases, Medely needed something low-lift, flexible, and seamless. The Challenge Medely evaluated a variety of feature platforms, and Chalk stood out for its ease of adoption, support quality, and self-serve design. The team spun up Chalk in just a few days, replacing the batch pipeline with a unified system where features are defined in Python and computed in real time. What Chalk delivered: 1. Self-serve infrastructure The manual infrastructure work disappeared. Medely engineers could modify feature logic, update schemas, and add data sources without touching Terraform or coordinating pipelines. Chalk handled the infrastructure layer automatically, functioning as a data engineering team delivered via software. A solution like Chalk is profoundly important to our team because it provided the ability to buy engineering talent. I can rely on the fact that it's self-serve. 1. Real-time computation Chalk directly queries Postgres at inference time, eliminating the 24-hour staleness problem. Medely’s pricing models now combine Snowflake historical aggregates with Postgres real-time signals, availability, facility demand, and current shift patterns—reacting to marketplace conditions as they happen. 1. Intuitive feature development Chalk's resolver architecture transformed how quickly Medely engineers could build features. Previously, building features from related data meant writing separate complex queries across multiple tables. To calculate anything from a professional's job history, they manually joined `professional` to `jobs` and computed the metric, a pattern repeated across dozens of features. In Chalk, the team defined the relationship once: a `professional` has many `jobs`. Resolvers could then chain off that relationship: `professional.booked_jobs` gave them the entire history. From completed counts to cancellation rates, features that would've required separate Snowflake queries now flow naturally from that single relationship. The features are just flowing so naturally ... I've never experienced anything like that, including my time at Spotify. What used to take dedicated queries and pipeline coordination now happens in minutes. The Solution From the first model Medely deployed with Chalk, the ROI was clear. The charge rate recommendation model is projected to generate $800K in annual net revenue through improved margins on job placements, enabled by real-time computation and self-serve infrastructure. The first product we ever deployed with Chalk paid for our team, probably more, in net revenue. Beyond the immediate revenue impact, Chalk transformed how quickly Medely's ML team could operate: - Feature freshness - Feature development - Experiment velocity - Data quality Metric - Daily batch features written to Redis (24-hour lag) - Separate Snowflake queries with manual joins + Terraform/pipeline coordination - Experiment cycles took ~2 months - BI, ML, and Product teams calculated features differently Before Chalk - Postgres queries at inference time - Define relationships once → features flow from simple resolver chains - Moving toward 2-week sprint cycles - Unified definitions across teams After Chalk Faster development and real-time features enable Medely to better match nurses and doctors to urgent shifts, filling critical healthcare needs more efficiently. Outcomes Medely's 2026 planning revealed how central Chalk had become. The team prioritized "Invest in Chalk" as a top initiative focused on unlocking more platform capabilities. Multiple models are now in experimentation for both pricing and matching use cases. With Chalk as its foundation, Medely is building toward: - Smarter professional-facility matching: Match medical professionals to shifts based on preferences and performance - Rates with administrator control: Human-in-the-loop pricing where facility managers fine-tune ML recommendations - Unified training and serving: Eliminating the current translation work by using Chalk for both model development and production Adopting Chalk is the biggest singular win I have had as an ML engineer at this company. Looking Ahead Feature changes required heavy infrastructure work Batch pipeline locked feature freshness at 24 hours Experiment cycles took ~2 months due to manual overhead GCP Healthcare Staffing Self-serve feature platform delivers data engineering as software Real-time computation reacts to live marketplace conditions Intuitive feature development unlocks 2-week sprint cycles RecSys (matchmaking), Dynamic pricing # Fintech Leader FIS Built a Real-Time Identity and Fraud Detection Platform Powered by Chalk source: https://chalk.ai/customers/fis Learn how FIS, a Fortune 500 payments leader, built a real-time fraud detection platform using Chalk. Compute hundreds of features in real-time, reduce fraud losses, and scale to billions of transactions. FIS, a Fortune 500 company and one of the largest payments processors in the world, built a production-grade fraud platform using Chalk’s compute-first architecture for feature engineering. The system powers real-time debit card authorization, computing hundreds of features per transaction in under 40ms within FIS’s secure environment. FIS is a Fortune 500 company and one of the largest payments processors in the world, powering nearly a quarter of all U.S. debit transactions. It powers core banking, payments, and risk management systems for thousands of financial institutions. Within FIS, the Financial Intelligence (Fintel) team was created to build next-generation identity and fraud detection systems that protect every transaction processed across debit, credit, and ACH. Debit card authorization sits at the center of global payments. It is one of the most complex, latency-sensitive, and high-stakes workloads in fintech. Every authorization must be evaluated in milliseconds, across billions of transactions, with zero margin for error. Our charter was simple but ambitious: to protect every transaction processed by FIS, across every bank and every rail. Overview The debit card authorization flow represents the holy grail of real-time decisioning. Fraud models must compute hundreds of behavioral features and return a decision instantly, at the moment a customer taps their card or completes a purchase. Legacy fraud platforms relied on static rules, models, and batch pipelines that could not adapt quickly to emerging fraud patterns or the scale of FIS’s transaction volume. The Fintel team set out to build a modern, ML-driven platform capable of operating at global scale while continuously learning from new behaviors and maintaining enterprise-grade reliability. They needed infrastructure that could meet demanding business, technical, and compliance requirements. Requirements: - Reduce fraud losses and false positives for issuers - Maintain frictionless customer approvals and minimize false declines - Compute hundreds of real-time features per transaction - Deliver sub-50ms latency at P99 to ensure instant authorization decisions - Scale across FIS’s global transaction volume, handling billions of events per day - Maintain compliance and data control We wanted to build the fraud platform of the future, one that uses ML to learn fraud patterns instead of relying only on static rules. The Challenge FIS selected Chalk to power real-time feature computation for its debit authorization models. Chalk’s compute-first architecture allowed FIS to define, validate, and serve features directly from existing data sources within their own environment and without data movement. The platform provided a unified framework for both online and offline feature computation, accelerating model development while maintaining compliance. - Unified feature computation - Built-in validation and testing - Materialized aggregates - Offline data acceleration - In-VPC deployment Capability - Enabled data scientists to define features once and reuse them across training and inference, ensuring consistency and reproducibility. - Ensured feature correctness through unit tests and eliminated online/offline skew. - Enabled real-time rolling computations, such as “transactions in the past 24 hours,” at sub-40ms P99 latency. - Generated training datasets in hours instead of weeks, improving experimentation velocity. - Ensured all data remained within FIS’s AWS environment, maintaining compliance protocols. FIS use case Chalk solved the online-offline skew problem completely. It even helped us diagnose where the skew was coming from. The Solution Chalk was deployed entirely within FIS’s AWS environment, running both the control and data planes inside the company’s EKS clusters. This architecture integrated Chalk’s compute-first design with FIS’s existing infrastructure, allowing the team to achieve high throughput while maintaining full data residency and compliance. - Performance - Scale - Data Sources - Models - Security Component - Consistent sub-40ms P99 latency for feature computation across hundreds of features per transaction. - Thousands of authorizations per second supported with minimal compute overhead. - Features computed directly from Snowflake and Databricks. - Multiple XGBoost models sharing consistent feature definitions - All data and computation remained within FIS’s VPC. Implementation We were computing hundreds of features for each transaction and still seeing P99 latency around 30 to 40 milliseconds. Architecture FIS went from proof of concept to production in just 12 weeks, meeting its goal of launching a fully operational, compliant, real-time authorization system faster than any prior internal effort. With Chalk, the Fintel team built a fraud detection engine that continuously learns from new data, scales globally, and supports rapid experimentation without adding infrastructure complexity. - Model iteration speed - Operational efficiency - Scalability Metric - Week-long data generation and training cycles - Manual offline feature generation and validation - Batch-oriented processing Before Chalk - 2 days or less iteration cycles, multiple models in parallel - Unified feature computation with automated validation - Real-time system supporting thousands of transactions per second After Chalk We went from sandbox to production in under twelve weeks. For a system of this complexity, that’s unheard of. Outcome FIS evaluated multiple feature store solutions, including Tecton and Fennel, but needed a system built for real-time performance and enterprise control. Chalk offered a compute-first approach to feature engineering, providing the functionality of a feature store with sub-40ms latency for hundreds of features and full deployment inside FIS’s environment. - Engineering partnership: Chalk’s engineering team partnered directly with FIS from pilot to production. - Collaborative support: Fast, hands-on communication replaced traditional vendor layers. Questions were resolved in minutes, not days. - Problem-solving mindset: Chalk adapted quickly to complex enterprise constraints, helping FIS meet aggressive goals without tradeoffs in compliance or scale. Other vendors sent sales engineers. Chalk sent the people who actually build the product. Why FIS Chose Chalk By adopting Chalk’s compute-first architecture, FIS proved that global-scale fraud prevention and enterprise compliance can coexist in real-time systems. The Fintel team runs a platform that adapts instantly to new fraud patterns, minimizes false positives, and delivers faster, more reliable decisions across billions of transactions each day. Our approach was to build the fraud platform of the future, and Chalk helped make that possible. Key Takeaway Compute hundreds of features per transaction in under 50ms at P99 latency for debit authorization Reduce fraud and false positives at massive global scale Maintain enterprise compliance while supporting billions of daily transactions AWS Finserv Compute-first features with sub-40ms at P99 latency Unified online + offline features with guaranteed correctness and consistency Ensured all data remained within FIS’s environment, maintaining compliance protocols Card Authorization # Powering Reliable, Explainable Data for Credit Decisioning at iwoca with Chalk source: https://chalk.ai/customers/iwoca Chalk powers iwoca’s credit decisioning infrastructure with precise, point-in-time feature computation. iwoca chose Chalk for our compute-first architecture, which unifies training and production pipelines, reduces drift, and makes data consistency a system-level guarantee. When a business seeks a loan, iwoca’s credit models ingest hundreds of features per application and produce credit decisions of up to £1,000,000 within seconds. There is no room for error; wrong inputs lead to wrong outputs or process failures. With hundreds of thousands of loan offers generated each year, a stable base for serving data is critical. For iwoca, accuracy and reliability are paramount. As one of Europe’s leading SME lenders, the company provides flexible financing to thousands of small and medium-sized businesses. Each decision must be consistent, defensible, and explainable. Data science is central to iwoca’s operations, their models underpin responsible lending and ensure every offer is based on precise and transparent data. Over time, complex ETL pipelines made maintaining that precision and generating training sets slower and more difficult to observe. Compared to some modeling problems, we make relatively few, very high-value decisions. Throughput isn’t the problem. Trust is. Overview After more than a decade of growth, iwoca’s feature pipelines had become fragmented. Training and production definitions diverged. Overnight ETL jobs were slow and unreliable, and teams spent more time managing pipelines and reconciling data drift than improving models. The problem was not data scale but data integrity. iwoca’s data is precise and time-dependent, combining bureau data, transaction records, and evolving cash-flow signals. Even a small timestamp error could change a credit score or repayment forecast. When assessing platforms like Tecton and Michelangelo, iwoca found they were designed for high-throughput workloads such as personalization. Chalk provided the same scalability with the temporal precision, traceability, and auditability required for financial modeling. We needed correctness — precise, point-in-time features for our data. The Challenge iwoca chose Chalk for its architecture, which aligned with the team’s engineering philosophy. Unlike traditional feature stores, Chalk computes features directly from source data and can persist them in online or offline stores when needed, but it does not rely on those stores as the source of truth. This architecture gives iwoca deterministic control over every feature. Features can be recomputed for any historical point in time from the source-of-truth data, ensuring accuracy without duplication or manual backfills. With Chalk, iwoca: - Generates point-in-time-correct training datasets using the same definitions as production - Recomputes features on demand for any historical period - Reduces drift between training and serving by keeping feature definitions aligned - Versions every dataset for complete traceability The result is a single system where consistency starts with training. Models trained on Chalk reflect exactly what was known at the time of decision, ensuring alignment between training and live prediction. It used to take 24 hours to regenerate a training set. Now it’s under an hour. We can trust every feature in it. Our data isn’t large, but every feature has to be precise. A single misaligned timestamp can change a lending decision. Chalk’s feature engine gives us conviction that every feature value is right. The Solution Traditional feature stores store precomputed features for offline and online systems, requiring constant synchronization. This ETL-driven approach causes drift, duplication, and uncertainty about which version of a feature is correct. We no longer maintain an offline store. Chalk computes training datasets directly from the data itself, exactly when we need them. Chalk’s feature engine reverses the ETL model by treating computation as the primary mechanism and storage as an optimization. Traditional systems treat stored features as the source of truth, while Chalk uses raw data as the authoritative source, resolving features through dependency graphs and recomputing them as needed. - Manually maintained offline store prone to missed or delayed updates - Separate pipelines for training and serving - Manual backfills and sync tasks - Offline store as source of truth Legacy Architecture - Features recomputed directly from source - Unified resolver graph governs both - Deterministic recomputation for any time point - Raw data as source of truth Chalk Feature Engine For iwoca, this model prevents inconsistencies that could alter credit outcomes and scales more reliably than synchronization-based systems. Instead of relying on stored data to stay in sync, we compute directly from the source. Chalk makes that both precise and reliable. By making computation the source of truth, Chalk gives iwoca a credit platform where data integrity is not maintained through process, but guaranteed by design. Architecture By rebuilding on Chalk, iwoca turned accuracy and explainability into system-level capabilities rather than operational goals. - Training data regeneration time - Feature definitions - Offline store - Data backfills - Confidence in inference-time features Metric - 24 hours - Separate for training and production - Required for every model - Manual and error-prone - Dependent on ETL consistency Before Chalk - Under 60 minutes - Unified resolver graph - Optional cache only - Deterministic, recomputed on demand - Guaranteed by architecture After Chalk Chalk reduced the time it takes to create training sets and improved our confidence in feature consistency. The training data is correct, the features are consistent, and that reliability matters more than anything. Today, iwoca’s credit decisioning runs on Chalk. Each loan, whether for a bakery, a manufacturer, or a seasonal retailer, is evaluated using features computed from the most current data available. The result is responsible, explainable lending at the pace of business. Outcomes Fintech and financial-services companies face the same challenge iwoca solved: unifying model training and serving in environments where correctness and consistency can outweigh raw speed. iwoca plans to expand Chalk’s use across credit decisioning, lifetime-value forecasting, and other areas of iwoca’s business. We don’t make millions of predictions a second. We make thousands of predictions of high consequence. Chalk underpins that work by giving us reliable, consistent data to power our models. Looking Ahead Fragmented feature logic between training and production pipelines Overnight ETL jobs that that eroded trust in recent data Timestamp drift affecting credit scores and auditability AWS Finance Unified resolver graph for consistent feature definitions across training and production Point-in-time recomputation from source-of-truth data Compute-first architecture with full versioning and auditability Credit Underwriting # MoneyLion delivers AI-powered personal finance products with Chalk source: https://chalk.ai/customers/moneylion MoneyLion uses Chalk’s real-time data platform to unify ML development, accelerate feature delivery, and reduce time-to-production. MoneyLion uses Chalk’s real-time data platform to unify ML development, accelerate feature delivery, and reduce time-to-production. By centralizing experimentation, serving, and governance, MoneyLion scales AI applications across fraud prevention, budgeting, and personalization. MoneyLion is on a mission to empower Americans to make better financial decisions. With millions of users across lending, investing, and personal finance tools, the company relies on a complex machine learning ecosystem to drive real-time fraud detection, customer engagement, and personalized recommendations. As the business scaled, building and deploying machine learning solutions across teams—ML operations, backend engineering, product, and data science—became more challenging. Each group owned a different part of the ML lifecycle, with distinct priorities: - Feature Platform Team (FPT) backend engineers optimized ingestion pipelines, minimized query latency, and ensured data quality. - MLOps engineers focused on scalability, governance, and platform reliability. - Data scientists needed to iterate quickly with Python, without infrastructure barriers. - Product managers pushed for reusable, production-grade features that could drive business results. The result was natural friction. Different workflows, goals, and metrics made collaboration slow and costly. MoneyLion needed more than a better feature platform. They needed an alignment layer for the entire ML lifecycle—from idea to production. That alignment didn’t mean forcing teams into the same process. It meant creating a shared space where data scientists could build, engineers could scale, MLOps could govern, and product teams could move quickly. Without that layer, features were delayed, duplicated, or lost between teams. Designing a feature platform that fits all team use cases—and still performs—was one of our biggest challenges. Overview MoneyLion’s first-generation feature platform was technically robust but operationally fragmented. Built around Java-based micro-services using SpringBoot, Postgres, and Redis, the system prioritized scalability — but also introduced complexity. Data scientists and ML scientists, like Jing, prototyped features offline in SQL or notebooks. Bringing these experiments into production required rewriting logic in Java, often by different engineers on the Feature Platform Team (FPT). This translation step added significant latency between experimentation and deployment. We would prototype offline, then wait for engineers to rewrite everything for production. It was slow and disjointed. The backend engineers, like Anya, maintained multiple custom micro-services for different products, with duplicated ingestion and feature computation logic. Without a centralized feature catalog or lineage tracking, feature reuse was rare and offline/online skew was common—especially for real-time applications. Despite heavy investment in Postgres query optimization and Redis caching layers to improve read performance for fraud use cases, maintaining sub-second latencies remained difficult under peak loads. Many fraud models required features to be computed and served within hundreds of milliseconds to meet strict SLA targets. There were no reusable features or central catalog. Every use case started from scratch, and supporting that at scale was overwhelming. MLOps engineers, like Melvin, managed governance manually — controlling data access, code promotion, and observability across a patchwork of microservices — without a unified interface for lifecycle management. Without a cohesive system, features were delayed, duplicated, or dropped—limiting MoneyLion’s ability to deliver real-time, AI-driven experiences at scale. The Challenge Chalk unified MoneyLion’s fragmented ML workflows by providing a developer-first platform that lets each team contribute effectively without forcing a rigid process. How MoneyLion transformed their ML workflow - Features prototyped offline, rewritten manually for production - No central feature store, catalog, or reuse - High engineering overhead for the FPT - Manual governance and slow approvals - Long delays from idea to deployment Before Chalk - Features built directly in Python and productionized with Chalk - Centralized catalog, automatic lineage, and easy feature reuse - FPT focuses on scaling, latency, and platform reliability - Built-in branching, isolation, and governance - Rapid iteration: hours/days vs. weeks After Chalk For backend engineers, Chalk replaced manual micro-service maintenance with dynamic feature pipelines. Engineers now focus on scaling system throughput, onboarding new real-time data sources, and optimizing query planners, rather than hand-translating feature logic. After switching to Chalk’s Query Planner, latency stabilized across peak load days. Even end-of-month Fridays, we stayed within SLA. For MLOps engineers, Chalk introduced a clean, branch-based development lifecycle. Each feature change exists in an isolated environment until promoted, reducing the risk of dependency conflicts or production regressions. Governance is enforced automatically through versioning, feature ownership tracking, and runtime policy checks. It’s easy to get started on Chalk, and the isolation model keeps everything safe by default. Teams don’t block each other anymore. For data scientists, Chalk eliminated the offline/online skew. Features are defined once in Python, immediately versioned, and served online through Chalk’s real-time query engine. This closed the loop between experimentation and production, dramatically speeding up iteration cycles. Chalk lets DS build in Python and Pandas—which is familiar—and not worry about infrastructure. We can experiment and iterate much faster. For product managers like Meng, Chalk unlocked better observability and feature reuse across lines of business, reducing duplicate effort and shortening time-to-market. Now teams contribute to features instead of requesting them. That changes everything. The Solution Chalk acts as the central feature platform between MoneyLion’s data infrastructure and real-time model serving systems. It abstracts feature computation, storage, and online serving behind a unified Python-first interface. Teams define, test, version, and serve features through Chalk while maintaining full traceability and runtime guarantees. Architecture Chalk helped MoneyLion not only unify the process of productionizing their ML but also accelerate customer-facing innovation across key areas. Today, with Chalk powering real-time ML infrastructure, MoneyLion can: - Protect users with real-time fraud detection: Features served within tight latency budgets enable faster risk decisions at the time of transaction. - Deliver smarter budgeting and spend insights: Personalized finance nudges are surfaced based on live transactional behavior. - Surface contextual feed and lifecycle recommendations: Adaptive recommendations dynamically adjust to user context in real-time. These capabilities are active today, reaching millions of MoneyLion users. Chalk helps us deliver financial products that are more responsive, more personalized, and more secure for millions of users. It’s a direct line from infrastructure to impact. Internal alignment drives external product impact dual-positive - Data engineers focus on platform optimization, not manual feature support - MLOps enforces platform governance automatically - Data scientists productionize features independently in Python - Product reuses features across lines Internal alignment - Higher availability of real-time features for fraud and finance - Faster and safer experimentation across teams - Smarter, faster model deployments - Shorter launch times for AI-driven features External product outcomes Outcomes Chalk now serves as the foundation for MoneyLion’s future AI strategy. As the company grows, Chalk is enabling: - Broader feature ownership across product lines. - Faster development cycles from experimentation to production. - Real-time architecture that scales with user growth and product and data complexity. We’re using Chalk to embed AI everywhere—from smart budgeting to fraud to lifecycle engagement. It’s foundational now. Looking ahead Fragmented collaboration across ML lifecycle High engineering overhead and offline-to-online drift Lack of centralized governance and feature reuse AWS Personal Finance Unified ML development on shared platform Python-native feature development for seamless experimentation-to-production Feature store with built-in governance and versioning Fraud, Recommendations # How Chalk fuels Whatnot’s live shopping platform source: https://chalk.ai/customers/whatnot Chalk powers Whatnot’s real-time recommendation engine, enabling low-latency personalization and scalable ML infrastructure for the largest live shopping marketplace. Whatnot uses Chalk to power real-time recommendations across its marketplace, from feed ranking to show discovery. By replacing a batch-based system with a unified feature platform, Whatnot delivers fresher, more personalized experiences, iterates on models faster, and scales its machine learning systems to meet the demands of its fast-growing live commerce platform. Whatnot is the largest live shopping platform in the U.S. and Europe, a marketplace where buyers and sellers connect through real-time, community-driven commerce. The company surpassed $3 billion in live sales last year alone, with users spending on average 80 minutes a day browsing and buying from live shows. As with most e-commerce marketplaces, personalization is critical. The first few seconds a user opens the app determine whether they’ll find a show they care about and whether a seller makes a sale. Feed recommendations are central to that moment. As Whatnot scaled, its ML systems started to lag behind product needs. The original batch-based architecture, which generated billions of user–show predictions nightly, struggled to keep up with fast-changing inventory, new seller launches, and real-time user behavior. Cold starts were common, coverage dropped, and model iteration slowed, all of which impacted the quality of the recommendations powering the core app experience. The Data & AI team began rebuilding the recommendation system to support real-time online inference. They needed infrastructure that could serve tens of thousands of features per request at low latency, across deeply nested graphs and high-throughput workloads. They chose Chalk to power all core recommendation systems, from feed ranking to show discovery. At Whatnot’s scale, we couldn’t deliver personalization without a real-time feature engine. Chalk is core infrastructure. Overview As a live e-commerce marketplace, Whatnot’s core ML task is matching buyers to sellers quickly and accurately. Each user session triggers a model that scores thousands of livestreams, many added moments earlier, and determines what to rank first. For much of the company’s growth, ranking and recommendations were powered by an offline pipeline that generated over 10 billion predictions each night. While effective at small scale, this architecture introduced sharp tradeoffs as traffic and seller volume grew: - New sellers were delayed: cold starts required 24 hours before personalized content appeared in ranked feeds - Session signals were dropped: the batch system couldn’t include recent behavior like taps or watches - Coverage fell: as inference scale grew, prediction coverage dropped to ~90% - Waste grew: trillions of scores were computed but never used - Model velocity stalled: new versions were hard to test and deploy bad All of this made it more challenging to iterate quickly, personalize effectively, and scale reliably. We needed a system that could reflect what was happening on the platform right now—not 12 hours ago. Moving to real-time inference was the only way to keep pace with our users and sellers. The Challenge Whatnot adopted Chalk to power its real-time feature infrastructure, serving as both the company’s feature store and online feature engine. The goal was to unify training and inference, reduce operational complexity, and enable low-latency inference without sacrificing flexibility or scale. Chalk serves as the system of record for features and as the real-time compute layer. Feature logic is defined once in code and reused across batch and online systems, reducing duplication and improving consistency. At inference time, Chalk materializes features from request context and historical data. Payloads exceed 1MB and include tens of thousands of features, yet the system consistently delivers responses under 150 milliseconds, even during peak traffic. We’re moving hundreds of millions of features per second, each payload around 1MB, and still hitting a P99 latency of just 100ms. That kind of performance across the board is a real testament to the system Chalk built. Chalk’s architecture allows Whatnot’s recommendation system to respond in real time to user behavior and marketplace activity, like browsing behavior, recent purchases, or seller activity. Because features are defined once and versioned in code, the team ships models faster and with more reliable deployment. By combining feature storage with on-demand computation, Chalk has become a foundational part of Whatnot’s ML platform. It supports production workloads while enabling fast experimentation across teams. Chalk gives us a single abstraction for online and offline features. That’s helped reduce bugs, speed up deployment, and made it easier to reason about feature correctness. The Solution Here’s how Chalk fits into Whatnot’s real-time inference loop: User opens app → request sent to backend Feed request triggered Backend retrieves live + upcoming shows Show inventory fetched Chalk computes real-time + historical features for each user–show pair Features computed Model scores all candidate shows based on Chalk features Model Scoring Ranked, personalized feed returned to user in <150ms Feed rendering System performance at a glance: - Supports 300M+ features/sec, including high-cardinality lookups and session-level joins - Low latency under load, with optimized paths ~70ms and P99 <100ms - >99.99% uptime - Load tested to handle 9x future traffic - Supports fresher, more powerful signals, including 1-hour lookbacks (moving toward 1-minute) Whatnot uses Chalk resolvers to compute features from data warehouses, streams, and request inputs — all through a unified query plan. The same feature definitions are reused across batch training and online inference, helping the team maintain consistency and move faster. Architecture With Chalk in place, Whatnot powers all feed recommendations using real-time online inference. The shift has driven measurable improvements across engineering efficiency and commercial performance: wideMetric - Cold start delay - Prediction latency - Dev iteration time - Personalization reach Metric - ~24 hours - Weeks - ~90% of users Before Chalk - <1hr delay - ~150ms end-to-end - Days - 99.9% of users receive real‑time personalized feed After Chalk Chalk worked with us on infrastructure tuning, latency tail issues, and scaling readiness. It isn’t just a tool — it feels like adding capability to our team. Outcomes What started with real-time feed ranking has steadily expanded. Today, Chalk supports a growing number of ML use cases across Whatnot’s business. - All core recommendation systems rely on Chalk to compute and serve real-time features for feed ranking, and show suggestions. - Fraud and trust models plan to leverage session-level and behavioral features, served in real time via Chalk without duplicating infrastructure. - Experimentation workflows use Chalk to snapshot feature values at inference time, ensuring test consistency and accurate evaluation. We’re constantly load testing to stay ahead of our own growth. Chalk has scaled with us from the early stages, and we’re confident it’ll support where we’re going. Chalk plays a foundational role in Whatnot’s growth. It provides stability under peak load, but more importantly, its composable design allows Whatnot to scale up model complexity without scaling infrastructure burden. By unifying, computing, and serving features in one system, Chalk helps ensure machine learning stays a technical moat — powering a fast, personalized experience as the marketplace grows. If you want to learn more about what Whatnot’s engineering team is up to, check out their blog. Looking ahead Coverage fell to ~90% as scale and inventory increased. Personalization was delayed up to 24 hours due to cold starts. Session-level signals like taps and chats were being omitted. AWS Marketplace Deliver low-latency computation at scale with 300M+ features/sec. Provide a unified feature store and real-time engine with Chalk. Enable faster deployment of models with fresher, live marketplace features. Recommendations # Mission Lane’s Source of Truth for Credit Decisioning with Chalk source: https://chalk.ai/customers/mission-lane Mission Lane uses Chalk to power real-time credit approvals, fraud detection and customer-facing features, all from a single, consistent feature platform. As modeling and decisioning became core to the business, Chalk gave the Mission Lane team the infrastructure to move faster, reduce drift, and scale confidently across systems. Mission Lane is a fintech helping millions of Americans left behind by traditional financial service companies access fair, transparent credit. Founded by industry veterans, the company has built a customer base of over 2.5 million by combining traditional credit data with machine learning. As modeling and decisioning systems became more complex—and more central to daily operations—Mission Lane needed a modern feature platform to unify real-time and batch infrastructure. They chose Chalk to serve as a centralized, production-grade feature store to power underwriting, fraud detection, and customer-facing product experiences. We're a fintech, so we are opportunistic in where we build our infrastructure, and we needed a partner who could solve this cleanly. Overview As Mission Lane’s modeling workflows matured, the team reached a familiar breaking point for many data-driven fintechs: feature engineering had become a bottleneck. The problem wasn’t model quality—it was the growing complexity of the systems around them. Across the company, different roles relied on different tools: Defining and implementing features in heterogeneous systems (including for production) Data scientists Building data ingestion pipelines Data engineers Implementing logic in production systems Engineers Each team operated with its own requirements, but without a shared foundation, teams rebuilt the same features multiple times across environments, leading to duplication, silent inconsistencies, and the potential for drift between training and production. This friction was amplified by the nature of the data itself: over 6,000 features spanning customer provided data, credit bureau pulls, Plaid transaction data, and multi-year customer behavior histories. These weren’t trivial aggregates—they required nested joins, temporal logic, and consistent semantics across systems. At the center of the problem was Mission Lane’s hybrid architecture. The company relied on both: - Batch workflows, like the Credit Line Increase Program (CLIP), which score hundreds of thousands of customers nightly - Real-time application decisioning, where features are computed in real time as a user submits an application In practice, this meant implementing the same feature logic in multiple places, across batch scoring, real-time inference, and model training pipelines, often with subtle differences that introduced inconsistencies. It wasn’t just about tech debt. You could train a model in Python, but when it went into production, it could get slightly different inputs. That’s a dangerous problem when you’re dealing with credit risk. To move faster and more safely, the Mission Lane team needed: - One place to define and calculate features and serve them across both batch and real-time systems - Support for Python and SQL, without requiring teams to choose between them - Reliable lineage and observability to satisfy engineering, risk, and compliance needs The Challenge Mission Lane chose Chalk as its next-generation feature platform to unify batch and online decisioning and eliminate drift across workflows. Part of what made Chalk the right fit was its ability to integrate cleanly into Mission Lane’s hybrid architecture and existing tools—Python, SQL, DBT, and Snowflake. Teams didn’t have to change how they worked; they simply plugged into a shared platform that standardized feature logic where it mattered. Chalk’s Kubernetes-native design and support for hybrid-cloud deployments also made it easy to run securely within Mission Lane’s own GCP environment, without introducing new infrastructure complexity. Chalk now powers both mission-critical model evaluations and emerging use cases beyond machine learning. Every feature defined in Chalk is automatically available across batch jobs, real-time inference, and reverse ETLs, with no need for duplicate engineering work. Chalk lets us define features once and use them everywhere—whether we’re evaluating two million customers overnight or scoring someone in real time as they apply. Use Case Description Credit Line Increase Program (CLIP) Evaluates 2.5M+ customers monthly to determine credit line increase eligibility and help customers grow financially with Mission Lane. Batch Live credit decisioning Real-time credit decisions and initial line assignments based on bureau and application data. Real-time Fraud detection Behavioral signals and payment patterns are evaluated during online interactions. Reverse ETL for credit score delivery Educational credit scores surfaced in-app via Chalk APIs, even for new users. Customer UX Agent tooling Live access to identity features for support agents, pulled from offline store. Ops What began as a solution for ML pipelines has become foundational to operations, support, and product. The Solution Mission Lane’s stack resembles many modern fintechs, but their hybrid workload model introduces real complexity. They needed infrastructure that could scale up for batch processing, scale down for low-latency online scoring, and integrate cleanly with their existing platform. Mission Lane’s modern data stack includes: - Data warehouse: Snowflake - Orchestration: DBT, Airflow - Model development: Python, Jupyter notebooks, DVC pipelines - Infrastructure: GCP + Kubernetes Chalk sits across both batch and online environments, supporting: - Unified feature definitions in Python and SQL - Consistent execution in batch and real-time environments - Reverse ETL APIs for production-grade CX and ops-facing applications - Auto-scaling compute for high-volume batch evaluations, like CLIP This hybrid architecture—where the same feature logic must support both a single API call and a batch job over millions of rows—is exactly where most tools break down. Chalk made the abstraction seamless. For us, real-time means computing features while the customer is waiting. Batch means aggregating behavior over time frames from days to years. We use Chalk for both, and it’s the same feature logic either way. Architecture Chalk helps Mission Lane move faster, improve reliability, and extend value beyond its ML team. wideMetric - Time to deploy new features - Training/serving consistency - Model iteration velocity - Reverse ETL access - Integration with new data Metric - Weeks - Prone to drift - Bottlenecked by engineering - Ad hoc, partial - Manual and slow Before Chalk - Days - Consistent across environments - Self-serve for data teams - Real-time and production-ready - Centralized and scalable After Chalk Chalk enables the next-generation rollout of CLIP—a core initiative affecting millions of customers. Outcomes Mission Lane set out to find a feature store to unify feature logic across batch and real-time ML. What they found in Chalk was a data platform with a feature engine and store built in, powering decisioning across risk, operations, and customer experience. Adoption continues to expand: - Faster integration of new bureaus and alternative data sources - Broader support for explainable credit models through transparent lineage - New customer features like credit score tracking over time - Wider use across fraud, CX, and product workflows Chalk solved the problems we brought it in to solve—and ended up helping us productionize far more of our data than we expected. Looking ahead Fragmented feature logic across Python, SQL, and production systems Inconsistent model behavior across batch, real-time, and training pipelines No centralized system to track, version, or reuse features across teams GCP Fintech Unified feature definitions across batch, real-time, and training pipelines Native support for Python, SQL, and hybrid cloud deployment Centralized platform for decisioning used by ML, ops, and product teams Credit underwriting, fraud detection, in-product data # How Apartment List uses Chalk to power personalized apartment recommendations source: https://chalk.ai/customers/apartment-list Apartment List uses Chalk’s real-time feature platform to power an individualized apartment search experience. Apartment List leverages Chalk’s real-time feature platform to deliver an individualized and intuitive apartment search experience. By unifying fresh data from SQL, streams, Python UDFs, Expressions, and APIs in real time, Apartment List runs sub-10ms queries for its dynamic recommendations as renters refine their search. Chalk's ability to ingest multiple data sources while maintaining ultra-low response times makes it an essential driver of Apartment List’s ML infrastructure. Apartment hunting is time-consuming and frustrating. Renters often sift through hundreds of listings, repeatedly adjusting price, location, and amenity filters, only to see results that don’t bring them closer to finding their next home. Unlike traditional listing sites, Apartment List redefines this experience by leveraging machine learning (ML) to deliver a curated, responsive, and highly personalized search. By dynamically updating recommendations based on renters’ preferences and in-session behavior, Apartment List ensures that users see the most relevant listings in real time. For Apartment List, delivering this level of real-time personalization requires integrating fresh data from SQL, streams, Python UDFs, Expressions, and APIs, all while maintaining sub 5ms response times. To ensure that renters instantly see up-to-date, high-quality listings, Apartment List turned to Chalk’s feature platform to compute in real time, accelerate ML deployment, eliminate data bottlenecks, and reduce latency across its search experience. Manually managing feature data worked—until it didn’t. As we scaled to multiple models, the cracks started showing. We needed a centralized system to handle features seamlessly and power real-time inference. Overview Finding an apartment is an interactive and personalized process. Renters continuously refine their budget, location preferences, and desired amenities—expecting search results to adjust instantly. If preferences fail to update in real time, users encounter stale or irrelevant listings, leading to frustration and drop-off. Ensuring that search results remained dynamic and highly relevant at scale became an increasing challenge. Apartment List pulled data from multiple sources, including databases, APIs, and internal services, introducing latency issues and making debugging complex. Without a centralized platform, the ML team struggled to maintain efficiency as its infrastructure expanded. Before implementing Chalk, Apartment List relied on custom-built feature-fetching pipelines, batch-processed data updates, and multiple disparate data sources to power its search personalization and ranking models. These manual processes created bottlenecks that slowed down search updates, particularly as the company scaled its ML-driven recommendations. As Apartment List’s ML capabilities grew, the team encountered three major pain points: - Managing real-time feature computation - Search results needed to update instantly when renters modified their preferences, whether adjusting their budget or refining location criteria. However, without a real-time feature platform, updates were delayed, preventing users from seeing the most relevant listings. - Scaling feature engineering across multiple ML models - Apartment List initially operated with a single ML model running in a bespoke Django-based system. As their ML expanded to power search ranking, price flexing, and location-based recommendations, the infrastructure became unsustainable. Engineers manually defined and fetched features for each new model, leading to duplicated logic, inconsistencies, and slow deployment cycles. - Latency and performance bottlenecks - Search ranking models that required querying multiple sources—including transactional SQL, real-time API calls, and streaming—introduced delays. Without inference-time compute, users engaged less with listings, impacting the overall experience. bad The Challenge: To overcome these challenges, Apartment List implemented Chalk’s feature platform, transforming how it delivers personalized, low-latency search recommendations. With Chalk, Apartment List: - Enables immediate search updates for a more responsive UI - ensuring search results adjust in real-time as renters modify their preferences. - Delivers dynamic personalization - ML models flex on price, location, and user behavior for more relevant recommendations. - Reduces latency for ranking - search rankings update in milliseconds, creating a faster, higher-performing experience. good Instant search updates: a more responsive UI With Chalk, search updates happen instantly. Renters no longer need to refresh or restart their search—listings update in real-time as they refine their criteria and perform more activity. This eliminates lag, improves engagement, and makes the apartment search feel fluid and interactive. Chalk’s performance directly affects the quality of our search and discovery models, which power everything from price flexibility to apartment ranking. The ability to call real-time features without dealing with stream complexity has been huge for us. Dynamic personalization: adapting to renter behavior Apartment List’s search experience doesn’t just show what users say they want—it adapts in real time to how they browse. With Chalk, ML models dynamically flex recommendations based on user behavior, making search results more intelligent and relevant. - Price Flexing: When a renter consistently interacts with listings above their stated budget, pricing models adjust price thresholds dynamically, surfacing more relevant options to reflect real user intent. - Geo-Spatial Flexing: Instead of static radius-based searches, Apartment List computes search relevance based on travel time, commute preferences, user mobility, and behavioral patterns, matching renters with listings that fit their lifestyles. By leveraging inference-time behavior signals, Chalk enables Apartment List to compute flexible, intent-driven search results so that renters see apartments they’re interested in, not just the ones they initially filtered for—without additional manual input. Chalk helped us move beyond static filters. Now, our search models adjust dynamically based on how renters actually interact with listings. Low-latency feature retrieval: faster ranking and real-time adjustments Powering a fast, ML-driven search experience requires more than real-time data ingestion—it requires a platform that can fetch, process, and deliver data in milliseconds. With Chalk, sub-5ms feature retrieval from multiple sources (SQL, APIs, and streaming data) ensures that search rankings remain dynamic and highly responsive. For the engineering team, the benefits extend beyond performance. Building new ML models is now exponentially faster. Before Chalk, launching a new ML model took weeks due to manual feature pipeline development. Now, features are defined once and reused across multiple models, reducing deployment time from weeks to days. Before Chalk, we had engineers writing custom feature-fetching services for every model—super time-consuming and brittle. Now, we can take a model from an endpoint to production in one to two days max, whereas before, it was a long, painful process. The Solution: Since implementing Chalk, Apartment List has significantly improved search personalization, system performance, and engineering efficiency: Chalk ensures search results adjust instantly when users change preferences Real-time search personalization Sub-5ms response times keep ranking updates highly performant Low-latency feature retrieval Chalk integrates with Apartment List's APIs, supporting both batch and streaming data services, unlike event-driven-only solutions. Seamless API integration Chalk works closely with Apartment List to optimize performance, resolve edge cases, and enhance ML infrastructure with custom features Strategic partnership I was relatively new to the MLOps space when I joined Apartment List, but the Chalk team worked closely with us to optimize performance, solve edge cases, and even build features we needed. Chalk feels more like a partnership than a vendor relationship. Outcomes With Chalk as the foundation of its ML infrastructure, Apartment List is focused on further refining its AI-driven personalization engine. The next phase includes experimenting and deploying new models, expanding real-time ranking optimizations as users engage with listings, incorporating even more dynamic user behavior insights, and further reducing latency to deliver a superior apartment search experience. By continuously iterating on its ML-powered recommendations, Apartment List is ensuring that renters find their perfect home faster, with a search experience that is intuitive, intelligent, and personal. Looking Ahead Delays in feature computation impacted search relevance for renters updating preferences Manually defining and fetching features led to duplicated logic, inconsistencies, and slow deployment cycles Querying multiple sources introduced data latency and performance bottlenecks GCP Marketplace Enables immediate search updates for a more responsive UI Delivers dynamic personalization for more relevant recommendations Supports seamless integration of heterogeneous data sources through direct API calls Recommendations # Verisoul stops fake accounts with Chalk’s real‑time feature platform source: https://chalk.ai/customers/verisoul Discover how Chalk’s platform helped Verisoul increase detection accuracy with Chalk’s powerful, efficient, and secure platform. Verisoul leverages Chalk’s real-time feature platform to combat evolving fake account threats, shipping detection updates 10x faster and using fresh inference-time data for 4x more accurate fraud detection. With complete auditability and reduced engineering overhead, Verisoul maintains speed, accuracy, and transparency—keeping critical platforms secure at scale. Fake accounts are proliferating at an unprecedented rate. With AI, bad actors can generate thousands of synthetic identities, bot-driven profiles, and coordinated fake accounts in minutes—exploiting promotions, distorting engagement metrics, and overwhelming platforms before detection can respond. Traditional detection tools often lag behind, allowing these malicious accounts to slip through unnoticed. The result? Increased hosting and marketing costs, fraudulent payouts, and financial losses from chargebacks. Chalk understood our problem from day one. Our first call was with a sales engineer who had deep technical knowledge—it felt clear and technically sound. Their docs were excellent and completely open. And they even got us up and running on GCP that weekend. Verisoul specializes in stopping the most advanced fake account bots before they cause harm. Serving as a centralized view for all fake account problems, Verisoul provides businesses with immediate, high-confidence decisions. Its platform analyzes device signals, behavioral patterns, biometric anomalies, and network-level risks in real time—stopping take-downs before they cause damage. But fake accounts are a moving target. Today’s threats include large-scale fake account creation, identity farming, and sophisticated multi-accounting techniques that evade traditional risk tools. Staying ahead requires more than accuracy—it demands speed. To meet this challenge, Verisoul turned to Chalk’s real-time feature platform. With Chalk, Verisoul can: - Ship detection updates 10x faster, reducing model iteration from days to hours. - Process risk signaling data that’s only available at inference time, removing complex ETL pipelines and stale data. - Ensure feature consistency and auditability, keeping fraud models explainable and reliable. good With Chalk, Verisoul stops fake account abuse, identity manipulation, and automated attacks before they happen—without compromising speed, accuracy, or transparency. Overview Verisoul’s fintech, gaming, SaaS, and market research customers face relentless attacks. The challenge isn’t just identifying fake accounts—it’s recognizing the evolving ways attacks manifest across industries: - Bot-driven fake account creation, flooding platforms with inauthentic users. - Location obfuscation through proxy networks and VPNs to bypass security. - Synthetic identities blending real and fake data to create seemingly legitimate profiles. - Multi-accounting and identity farming manipulating promotions, reviews, and credit systems. - Fraudulent engagement and data pollution distorting analytics and inflating KPIs. These threats don’t just pose financial risks—they erode user trust, damage platform integrity, and create long-term reputational harm. Before implementing Chalk, Verisoul relied on manual, in-house feature engineering pipelines, requiring Python notebooks and BigQuery queries to be manually converted into API logic. This led to three major bottlenecks: - Data staleness – Data for decisions needs to be fresh, but the existing infrastructure, built for batch or streaming data, couldn’t support real-time feature computation. - Slow signal iteration – Managing multiple feature pipelines across research, training, and production environments made testing and deploying new models slow and inefficient. Even minor updates required full dataset recomputation. - Lack of auditability – Customers need transparency on fraud decisions, but engineering teams had to manually pull and piece together computed feature tables, slowing response times and debugging efforts. bad Challenges Decisions Based on Inference-Time Data By leveraging Chalk, Verisoul has transformed its decision models into a true real-time system. Instead of relying on outdated or precomputed risk signals, they now make decisions using fresh, continuously updated data—ensuring the most accurate fraud detection possible. Even a 100-millisecond delay in data freshness affects fraud detection accuracy… Real-time processing makes us up to 4x better than traditional IP-based fraud lists. Instant Fraud Detection Iteration Risk detection requires analyzing vast amounts of nuanced signals, from behavioral patterns to network data. Chalk brings all these elements together in a unified system, enabling Verisoul to easily manage and refine their models. Instead of juggling fragmented datasets and manual updates, their team can now iterate on detection strategies within a single, centralized framework—enhancing accuracy and scalability. Verisoul’s engineers can now instantly recompute and validate new fraud signals using unified feature pipelines, allowing them to deploy updates 10x faster and stay ahead of evolving threats. Our product provides a decision—fake, suspicious, or real—but behind that, we analyze multiple sub-scores and thousands of features. Chalk brings all these pieces together in one place, allowing us to easily recompute and refine them, which is massively valuable. Balancing Speed with Auditability With Chalk, Verisoul gained full transparency into every signal. Every decision is now tied to auditable feature logs, making it painless to provide customers with clear, actionable insights into why an account was flagged. We had partially computed features saved in tables that took time and effort to pull for customers—now we have auditability for every signal in real-time with Chalk. Solutions Since implementing Chalk, Verisoul has significantly improved detection speed, accuracy, and scalability. Verisoul now delivers risk scores using the freshest data, making detection 4x more accurate than traditional rule‑based methods. Real-time fake account prevention at scale With Chalk, Verisoul’s engineering team can iterate on fraud signals without infrastructure bottlenecks, deploying updates within hours instead of days. 10x faster feature development and deployment Engineers no longer need to maintain separate feature pipelines. Fake account decisions are now fully traceable and explainable, strengthening customer trust. Full auditability & reduced engineering overhead Verisoul’s team now writes less code while shipping more powerful models—freeing up engineering resources for higher-value initiatives. Increased engineering efficiency and speed Outcomes Fake account tactics evolve daily—Verisoul moves faster. With Chalk as its feature engineering backbone, Verisoul has built an industry-leading real-time fake account detection engine that processes thousands of decisions per second using inference-time data. The next phase? Verisoul is working on introducing a new paradigm into fraud detection—enabling every customer to customize the intelligence and decisioning to their specific definition of a fake user. With Chalk, Verisoul can dynamically fine tune the underlying data and models to optimize the accuracy for each customer’s use case. By combining speed, precision, and adaptability, Verisoul continues to set the standard for real-time fake account prevention. Our competitive edge hinges on the speed at which we analyze risk patterns, test new detection features, and deploy updates. With Chalk, our iteration speed went from days to hours. Looking Ahead Data staleness for risk decisions impacted decisioning speed and quality Slow and inefficient signal iteration causing longer deployment Lack of auditability created engineering and customer service pain GCP Fraud Detection Now deploys detection updates 10x faster Real-time inference data improves detection accuracy 4x over traditional methods Full auditability enables transparent and explainable fraud decisions for customers Detection Models: risk tree decisioning # Vital predicts hospital wait times with Chalk’s feature platform source: https://chalk.ai/customers/vital Discover how Vital predicts hospital wait times with Chalk’s powerful, secure, and reliable feature platform. Using advanced AI, Vital transforms complex health record data into easy-to-use, personalized interfaces that inform and engage over one million patients per year. After deploying Chalk, Vital was able to stop wrangling infrastructure and focus on improving its models, launching new products, and delivering a world-class patient experience. Vital is redefining patient experience with software that gives more control, clarity, and predictability to emergency department visits and hospital stays. Using advanced AI, Vital transforms complex health record data into easy-to-use, personalized interfaces that inform and engage over one million patients per year. Hospitals across the U.S. use Vital to improve patient satisfaction, drive growth and patient loyalty, achieve better clinical outcomes, and reduce workload for care teams. However, Vital found that its previous solution’s data infrastructure was not up to the task —they struggled to make updates to their AI models, and spent more time managing data pipelines than innovating on their core product. After deploying Chalk, Vital was able to stop wrangling infrastructure and focus on improving its models, launching new products, and delivering a world-class patient experience. We probably write about half or a third of the code that we did with our previous solution — and get twice as much done. It’s a magical developer experience. The code is also much simpler. We’ve had developers from other teams come in and immediately understand stuff, which was 100% not the case before. Do more with less code. Overview Vital began their machine learning journey with another managed feature platform which ultimately did not serve their needs. Heavily dependent on Spark and Databricks, the previous solution created challenges for Vital due to its poor developer experience and architectural limitations: - Experimentation was a slow, tedious process filled with guess work and unreliable results, which made them unable to re-train their models. - Changes to feature pipelines required deep knowledge of Spark and Databricks in addition to the feature store architecture, which was limited to a small number of engineers. - Data processing occurred on third-party infrastructure. bad Challenges Vital selected Chalk to address each of these challenges. With Chalk, Vital dramatically increased the pace of its product development. - Rapid iteration and experimentation: Vital now deploys model updates 2-4 times per month. - Vital has up to nine engineers working concurrently on Chalk without being blocked on infrastructure. - Full control of customer data: patient data never leaves Vital’s cloud infrastructure. good Solutions One of Vital’s early models was trained on hospital data collected during COVID-19. Unsurprisingly, hospital activity during the early pandemic was vastly different from hospital activity post-COVID-19 vaccine, so it was crucial to re-train models to deliver accurate predictions. Vital recognized the importance of updating their models, but they were unable to release updates with the feature engineering solution they had. There were several reasons why model releases were challenging. To start, Vital’s engineers needed to manage Spark, Databricks, various custom data pipelines, and the feature store itself. Modifying a single feature required full re-computation of the entire feature view, so data scientists needed to coordinate with infrastructure engineers throughout experimentation. Even after data scientists were ready to bring an idea to production, they were unable to reuse training feature code for production feature generation, leading to training-serving skew. Re-implementing feature definitions led to inconsistency between training and production, which made it difficult for the team to be confident in the accuracy of production results. Chalk solved these problems. It unlocked rapid iteration cycles and empowered a broad range of engineers to contribute code. Critically, Chalk enabled users to reuse the same feature code between notebooks, training, and serving. Today, when Vital engineers want to experiment with new feature pipelines, they create branch deployments with a single Chalk command. Later, they can deploy the exact same code to production. Engineers no longer have to manually update Spark and Databricks — updating features is as simple as editing a few lines of code. “Anyone who can write Python can work with Chalk,” says Mack Delany, Director of Machine Learning at Vital. With Chalk, we can quickly add or remove a feature, test across all models, and then roll out quite quickly. For a team of nine that covers research and engineering for three different product lines, being able to iterate quickly and get improvements out and then move back to our roadmap was like the big shining light. Rapid iteration and experimentation Vital handles sensitive healthcare information, which is regulated by HIPAA and requires strict data control. Their previous solution required data processing to occur on third-party cloud infrastructure. When they set out to replace their machine learning partner, they searched for an enterprise-grade solution that could guarantee data would never leave their own cloud infrastructure. Chalk was a natural fit because it was designed with data privacy and security in mind from day one. - Chalk is deployed in Vital’s own cloud infrastructure and no production traffic exits Vital’s virtual private cloud (VPC). - Real-time data processing, feature caching, and query serving all happen from within Vital’s infrastructure. By deploying Chalk into a customer’s own cloud infrastructure, customers are able to retain control of their data while still benefiting from Chalk’s best-in-class developer experience. Full control of customer data Vital now uses Chalk to process and serve features powering seven different models in production. These models are used in real-time applications, such as Vital’s ERAdvisor software, which guides patients through emergency room visits with personalized wait times and next steps. ERAdvisor uses Chalk to provide accurate wait time estimates based on real-time, hospital-specific factors, including each hospital’s latest wait time data. Vital has seen improvements in: The team now deploys updated models 2-4 times a month. Execution Speed Vital is able to deploy model updates with confidence. Model accuracy Vital is able to serve upwards of 200 model predictions per patient visit with millisecond latency and low costs by leveraging Chalk’s real-time feature resolution and caching. Performance & cost efficiency All of Vital’s data storage, feature processing, and serving happens in its own cloud infrastructure, ensuring sensitive patient and healthcare data never leaves its environment. Data privacy and security Vital’s machine learning team has grown to support 9 engineers concurrently working on changes. Vital engineers from across the company are genuinely excited when they get to work on the Chalk codebase. Developer happiness Outcomes Vital is looking forward to expanding its product offerings with machine learning powered by Chalk. They aim to continue guiding patients through their healthcare journeys even after the emergency room. Vital’s latest offering, AccessAdvisor, helps patients connect with specialized doctors for follow-up care while considering insurance matching, proximity, and relevant experience. Vital will use Chalk to predict the best doctor matches for patients. Additionally, they plan to increase the granularity of their existing guided products, with new predictions such as wait time estimates for lab results. Looking ahead Experimentation slow, tedious, unreliable Feature changes required deep, specialized knowledge limited to a handful of engineers Third-party data infrastructure AWS Healthcare Now update model 2-4 times / month Up to nine engineers working concurrently Patient data never leaves the cloud ERAdvisor: Real-time hospital wait times # Run Chalk in your cloud source: https://chalk.ai/deployments/self-hosted Self-host Chalk's feature store on AWS, GCP, or Azure. Meet data residency and regulatory requirements without sacrificing real-time ML performance. TALK TO AN ENGINEER SELF-HOSTED Deploy Chalk directly into your AWS, GCP, or Azure environment. Keep your data, features, and models within your VPC while leveraging Chalk’s feature store. Trusted by teams building the next generation of AI + ML Chalk separates between the data and control plane. The data plane runs entirely inside your cloud. Feature computation, retrieval, and model inputs execute within your VPC. Your data stays in your environment. The control plane handles orchestration, configuration, and metadata. It can run inside your infrastructure or in Chalk’s cloud. Built for security-first teams Deploy Chalk directly inside your AWS, GCP, or Azure environment. Keep all data, features, and models within your VPC while operating under your own IAM, networking, and encryption standards. Control without compromise Chalk’s engine runs on Kubernetes in your cloud environment as a gRPC service. Feature computation and retrieval execute in single-digit milliseconds, even as workloads scale horizontally across models and teams. Performance at scale Run Chalk on your existing infrastructure to make use of your cloud credits and commitments. Simplify billing and reduce operational costs while maintaining real-time performance. Optimize cloud costs Full control. Production-grade performance. All data, features, and models remain inside your infrastructure. For our high-funnel search workloads, predictability matters. - Data control and residency - Security and network control - Performance and scale - Control plane - Infrastructure access - Observability - Maintenance and updates Choose your deployment option - All data, features, and models stay in your infrastructure. Region control aligns with governance and regulatory needs. - Inherits your IAM, networking, encryption, logging, and key management policies. Supports fully air-gapped deployments. - Single-digit millisecond local feature retrieval using Chalk’s runtime, scaling on your infrastructure. - Control plane runs in your cloud or Chalk’s, managing orchestration, configuration, and metadata. - Easily configure access to data sources in your own cloud, where it already lives. - Export metrics and traces to Datadog, Prometheus, NewRelic, and more. - Managed by your team, with optional Chalk-assisted upgrades. Self-hosted - Data processed in a Chalk-managed, tenant-isolated VPC with multi-region support. - Managed by Chalk with fine grained access controls and SOC 2 Type II, ISO 27001, and GDPR-compliance. - Single digit millisecond feature retrieval with auto scaling on Chalk managed compute clusters. - Fully managed orchestration layer maintained by Chalk. - Chalk assumes a fixed role you assign in your cloud. Requires networking setup. - Fully managed by Chalk with continuous updates. Chalk-hosted /deployments/self-hosted Run Chalk without managing infrastructure Chalk Cloud delivers real-time ML infrastructure as a service that is optimized, secure, and ready for production. Explore more of Chalk’s data platform # Run Chalk as a fully managed platform source: https://chalk.ai/deployments/chalk-hosted Chalk-hosted feature store with enterprise-grade performance. Single-digit millisecond retrieval, built-in observability, and SOC 2 / ISO 27001 compliance. No infra required. TALK TO AN ENGINEER Chalk-hosted Ship features in real-time without provisioning clusters, configuring orchestration layers, or managing scaling policies. Chalk operates the infrastructure so your team can focus on your models, features, and shipping value faster. Trusted by teams building the next generation of AI + ML Deploy instantly with Chalk’s fully managed environment. Your team can compute, store, and serve real-time features without provisioning clusters or maintaining pipelines. Focus on innovation, not infrastructure Chalk’s engine runs on Kubernetes in your cloud environment as a gRPC service. Feature computation and retrieval execute in single-digit milliseconds, even as workloads scale horizontally across models and teams. Performance at scale Monitor every feature, dependency, and service in real time. Chalk provides feature level metrics, request tracing, dependency graphs, and runtime health dashboards. Observability built in Enterprise-grade performance. Real-time performance and scalability without moving data. We’re applying AI and ML at scale across key areas of our energy business with Chalk’s feature platform. It enables high-performance computation over diverse data sources using clean, reusable code. - Data control and residency - Security and network control - Performance and scale - Control plane - Infrastructure access - Observability - Maintenance and updates Choose the model that fits your operational needs - Data processed in a Chalk-managed, tenant-isolated VPC with multi-region support. - Managed by Chalk with fine grained access controls and SOC 2 Type II, ISO 27001, and GDPR-compliance. - Single digit millisecond feature retrieval with auto scaling on Chalk managed compute clusters. - Fully managed orchestration layer maintained by Chalk. - Chalk assumes a fixed role you assign in your cloud. Requires networking setup. - Export metrics and traces to Datadog, Prometheus, NewRelic, and more. - Fully managed by Chalk with continuous updates. - All data, features, and models stay in your infrastructure. Region control aligns with governance and regulatory needs. - Inherits your IAM, networking, encryption, logging, and key management policies. Supports fully air-gapped deployments. - Single-digit millisecond local feature retrieval using Chalk’s runtime, scaling on your infrastructure. - Control plane runs in your cloud or Chalk’s, managing orchestration, configuration, and metadata. - Easily configure access to data sources in your own cloud, where it already lives. - Managed by your team, with optional Chalk-assisted upgrades. Self-hosted /deployments/self-hosted Talk to an engineer Run Chalk without managing infrastructure Chalk Cloud delivers real-time ML infrastructure as a service that is optimized, secure, and ready for production. Explore more of Chalk's data platform # Chalk at The AI Conference 2026 source: https://chalk.ai/events/AI-Conference-2026 Meet us in San Francisco and see how Chalk provides the data and infrastructure solutions to deliver real-time context to models and agents, train and fine-tune LLMs, and evaluate agents in production. From sub-5ms feature serving to an enterprise-grade agent runtime, build with the full AI/ML data platform for inference - entirely inside your cloud. This event has passed. Meet us at The AI Conference 2026 Schedule Meeting Speaking Sessions Meet us in San Francisco. See how Chalk provides the data and infrastructure solutions to deliver real-time context to models and agents, fine-tune LLMs, and evaluate agents in production. From sub-5ms feature serving to an enterprise-grade agent runtime, build with the full AI/ML data platform for inference - entirely inside your cloud. A premier conference for builders, researchers, and AI leaders building what’s next. September 29 – October 1 2026 When Pier 48, Mission Rock San Francisco, CA 94158 Where Shaping the future of AI Watch Chalk work its magic live — no slides, just real features doing real things. Live product demos Skip the pitch deck. Pull up a chair and geek out with our forward deployed engineers, one-on-one. 1:1 conversations See how top AI and ML teams are using Chalk to build agents, deploy SOTA models, detect fraud, authorize payments and much more, all in real-time. Featured use cases Fresh drops, limited supply, zero guarantees — grab your Chalk gear before it's gone. Latest swag™ Visit us at booth #262 Stop by our booth to see live demos, meet our team and learn how Chalk powers real-time AI applications. Stop by the Chalk booth (#262) at Pier 48 to see live demos, meet our team and learn how Chalk powers real-time AI applications. Moscone Center, San Francisco, CA, Booth #2611 Chalk Booth Visit Add booth visit to schedule For many teams, the promise of AI agents has been falling flat. Building agents has gotten easier, but productizing has become a whole other challenge. What’s often missed is that many agent failures that look like model or prompt problems are actually evaluation problems. The agent made a confident decision on data that was seconds, minutes, or even days out of date, and it’s hard to validate why it made certain decisions. In this talk, Chalk Co-Founder Elliot Marx will break down how inference-time data and infrastructure have become the critical layers every AI agent depends on, regardless of which model or orchestration framework it sits on. Attendees will walk away with a clear understanding of where batch-first stacks break under real-time demands, how the shared context layer improves models, and how best to combine data and infrastructure that enables AI agents to thrive in production. Pier 48, Theater 2 October 1, 2026 / 1:00 PM - 1:25 PM Chalk Technical Deep Dive SAVE TO CALENDAR Most AI failures that look like model or prompt problems are actually context problems - agents deciding confidently on stale data, and evals that can't predict production performance. Chalk Co-Founder Elliot Marx shows how a unified context layer fixes both: live, governed data served at inference time, queried where it already lives, powering evals with the same system. He'll demo an end‑to‑end workflow for shipping agents with confidence on this context layer, by simulating historical scenarios by giving the agent time-bound data and tools, evaluating on tasks, and deploying in a sandbox. Finally, Chalk will share lessons from building context infrastructure for teams like Whatnot, Turo, and Mission Lane Technical Deep Dive Two Halves of Inference to Build, Eval & Run Agents Speaking sessions Join us for insightful sessions featuring Chalk experts sharing real-world experiences and technical best practices. Meetings & conversations Connect with Chalk's executive team, techincal leaders and subject matter experts at The AI Conference 2026 to see how real-time feature pipelines power smarter ML. REQUEST A MEETING terminal Technical meetings Deep-dive sessions with Chalk engineers — architecture reviews, integration workshops, proof-of-concept discussions. user Executive meetings Strategic conversations with Chalk leadership — roadmap, partnership, and go-to-market alignment. Learn more about who we are, what we do and why it matters for teams operationalizing AI at real-time scale. Chalk 101 See how Turo uses Chalk's real-time data platform to standardize feature delivery, speed up iteration, and scale real-time ML across search, pricing, and risk. Customer story Chalk was recognized as one of the world’s most innovative companies of 2026 by Fast Company. World's Most Innovative Discover how Chalk helps teams build real-time models for mission critical operations. See Chalk in action Explore more of Chalk Meet the team # Chalk at Fintech Devcon 2026 source: https://chalk.ai/events/fintech-devcon-2026 Meet us in Denver and see how Chalk powers real-time feature engineering to make AI real for fintech builders shipping models and agents. This event has passed. Meet us at FinTech DevCon Schedule Meeting Speaking Sessions Fintech_devcon is a conference designed to educate and empower fintech builders. They'll learn hands-on tools, best practices, and industry secrets from actual builders in fintech today. August 3 - 5 2026 When Sheraton Denver Downtown Hotel Denver, CO Where The premier conference for fintech developers Watch Chalk work its magic live — no slides, just real features doing real things. Live product demos Skip the pitch deck. Pull up a chair and geek out with our forward deployed engineers, one-on-one. 1:1 conversations See how top ML teams are using Chalk to build recommender systems, detect fraud, authorize payments and much more, all in real-time. Featured use cases Show up every day for a shot at winning. Luck favors the ones who keep coming back. Daily proof solves Fresh drops, limited supply, zero guarantees — grab your Chalk gear before it's gone. Latest swag™ Visit us Stop by our booth in the South Convention Lobby to see live demos, meet our team and learn how Chalk powers real-time AI applications. Stop by the chalk booth in Moscone South to see live demos, meet our team and learn how Chalk powers real-time AI applications. Moscone Center, San Francisco, CA, Booth #2611 Chalk Booth Visit Add booth visit to schedule Sheraton Denver Downtown Hotel, Tower D Wed, Aug 5 · 2:00 PM MDT – 2:45 PM MDT Chalk Technical Deep Dive Model development is often slowed down by long iteration cycles, too many tools, and frequent handoffs. We learned this the hard way working with an enterprise building real-time fraud detection models for card authorization. This talk tells the story of building an AI agent that helps data teams investigate missed fraud, analyze model behavior, and automatically propose new rules and features. We realized early on that the agent needed access to the same context as the model, but the infrastructure underneath couldn't support that. The core of the session walks through rebuilding the foundation around a real-time context layer that computes fresh data from the source at inference time. We’ll highlight the limitations of the original batch-first systems: stale features, train/serve skew, inconsistent feature definitions, and latency constraints. Then, we'll discuss key agent and model engineering decisions, including navigating deployment model constraints, managing tradeoffs between freshness and speed, and treating observability as a requirement. Attendees will leave with a clear understanding of where batch-first stacks break under real-time demands, how the shared context layer improves models, and how AI agents can turn model development into a continuous cycle. Technical Deep Dive Rewiring the Fraud ML Workflow: How a Context Layer and an AI Agent Put Better Models in Production America/Denver Speaking sessions Join us for insightful sessions featuring Chalk experts sharing real-world experiences and technical best practices. Meetings & conversations Connect with Chalk's executive team, technical leaders and subject matter experts at Fintech Devcon to see how real-time feature pipelines power smarter ML. Schedule a Meeting with Chalk SCHEDULE A MEETING terminal Technical meetings Deep-dive sessions with Chalk engineers — architecture reviews, integration workshops, proof-of-concept discussions. user Executive meetings Strategic conversations with Chalk leadership — roadmap, partnership, and go-to-market alignment. Learn more about who we are, what we do and why it matters for teams operationalizing AI at real-time scale. Chalk 101 See how MoneyLion uses Chalk’s real-time data platform to unify ML development, accelerate feature delivery, and reduce time-to-production. Customer story Chalk was recognized as one of the world’s most innovative companies of 2026 by Fast Company. World's Most Innovative Discover how Chalk helps teams build real-time models for mission critical operations. See Chalk in action Explore more of Chalk Meet the team # Chalk at AI Council 2026 source: https://chalk.ai/events/ai-council-2026 Meet us in San Francisco and see how Chalk powers real-time feature engineering to make AI real for businesses. This event has passed. Meet us at AI Council Schedule Meeting May 12-14, 2026 / San Francisco, CA Meet us in San Francisco and see how Chalk powers real-time feature engineering to fulfill the promise of ML models for an AI world. Watch Chalk work its magic live — no slides, just real features doing real things. Live product demos Skip the pitch deck. Pull up a chair and geek out with our forward deployed engineers, one-on-one. 1:1 conversations See how top ML teams are using Chalk to build recommender systems, detect fraud, authorize payments and much more, all in real-time. Featured use cases Show up every day for a shot at winning. Luck favors the ones who keep coming back. Daily proof solves Fresh drops, limited supply, zero guarantees — grab your Chalk gear before it's gone. Latest swag™ Keep your eyes peeled for the fastest moving ride in town - and no, it's not just our ultra-fast data pipelines. Buckle up Visit us in the exhibitor hall Stop by our booth in the exhibitor hall to see live demos, meet our team and learn how Chalk powers real-time AI applications. Stop by the chalk booth in Moscone South to see live demos, meet our team and learn how Chalk powers real-time AI applications. Moscone Center, San Francisco, CA, Booth #2611 Chalk Booth Visit Add booth visit to schedule Meetings & conversations Connect with Chalk's executive team, techincal leader and subject matter experts at AI Council to see how real-time feature pipelines power smarter ML. Schedule a Meeting with Chalk SCHEDULE A MEETING terminal Technical meetings Deep-dive sessions with Chalk engineers — architecture reviews, integration workshops, proof-of-concept discussions. user Executive meetings Strategic conversations with Chalk leadership — roadmap, partnership, and go-to-market alignment. Learn more about who we are, what we do and why it matters for teams operationalizing AI at real-time scale. Chalk 101 See how Whatnot uses Chalk to power real-time recommendations across its marketplace. Customer story Chalk was recently recognized as one of the world’s most innovative companies of 2026 by Fast Company. World's Most Innovative Discover how Chalk helps teams build real-time models for mission critical operations. See Chalk in action Explore more of Chalk Meet the team Marc spent many years at Google where he helped to launch the first version of Google Wallet. He went on to start Index, which Stripe acquired as its in-store payment solution—now called Stripe Terminal. Elliot started his career at Affirm where he built the early risk and credit data infrastructure system (the inspiration for Chalk). He then co-founded Haven Money, which Credit Karma acquired to power its banking products. Andy worked at Palantir on large government data infrastructure projects. He then co-founded Haven Money (with Elliot), which now powers Credit Karma Money. Alex is a seasoned enterprise SaaS GTM leader, scaling marketing and revenue teams from Series A through C across categories including AI, HRTech, and security. Her expertise has helped build beloved brands and turn pipeline engines into category-defining growth machines. Melanie manages the post-sales forward deployed engineering and customer success teams at Chalk. Melanie previously worked on data infrastructure at Plaid, in addition to time at Airbnb, Two Sigma, and several founding teams. Samuel works on the forward deployed engineering team, helping customers get the most out of Chalk. Previously, they built and deployed ML models at Pattern Ag, a biotech company, and worked at Algolia. They studied Computer Science and English at Stanford. Collin works on the forward deployed engineering team, helping customers utilize Chalk to its fullest potential. He previously worked at Capital One across data lineage and model development-focused teams, as well as player retention and game mode development in the indie gaming industry. Rishi handles developer relations at Chalk, where he does everything marketing and engineering. Rishi previously worked on payments as a software engineer at Coinbase, and ML infrastructure as a product manager at Chime. # Chalk at Snowflake Summit 2026 source: https://chalk.ai/events/snowflake-summit-2026 Meet us in San Francisco and see how Chalk powers real-time feature engineering to make AI real for businesses operating with the Snowflake AI Data Cloud. This event has passed. Meet us at Snowflake Summit Schedule Meeting RSVP for Events Speaking Sessions Snowflake Summit is the premier annual conference for the AI data cloud, where thousands of professionals, developers, and executives gather to explore AI, machine learning, and app development. June 1-4 2026 When Moscone Center San Francisco, CA Where Booth #2611 Chalk Experience the future of enterprise data and intelligence Watch Chalk work its magic live — no slides, just real features doing real things. Live product demos Skip the pitch deck. Pull up a chair and geek out with our forward deployed engineers, one-on-one. 1:1 conversations See how top ML teams are using Chalk to build recommender systems, detect fraud, authorize payments and much more, all in real-time. Featured use cases Show up every day for a shot at winning. Luck favors the ones who keep coming back. Daily proof solves Fresh drops, limited supply, zero guarantees — grab your Chalk gear before it's gone. Latest swag™ Keep your eyes peeled for the fastest moving ride in town - and no, it's not just our ultra-fast data pipelines. Buckle up Visit us at booth #2611 Stop by our booth in Moscone South to see live demos, meet our team and learn how Chalk powers real-time AI applications. Stop by the chalk booth in Moscone South to see live demos, meet our team and learn how Chalk powers real-time AI applications. Moscone Center, San Francisco, CA, Booth #2611 Chalk Booth Visit Add booth visit to schedule In this session, Grindr's CPO, AJ Balance, and Chalk's CEO, Marc Freed-Finnegan, break down the Snowflake and Chalk stack that’s powering Grindr’s AI and what it means to ship real AI products at consumer scale for a global community of tens of millions. Monday, June 1 @ 1:00 PM Grindr x Chalk Customer Story The Infrastructure Behind Grindr’s AI Transformation In this talk, Chalk's Co-Founder, Elliot Marx, unveils a computation-first feature store architecture where Chalk extends Snowflake into real-time decision systems. Tuesday, June 2 @ 01:00 PM Chalk Technical Deep Dive Technical Deep Dive A Computation-First Architecture for Real-Time ML Feature Stores Speaking sessions Join us for insightful sessions featuring Chalk experts and customers sharing real-world experiences and technical best practices. Meetings & conversations Connect with Chalk's executive team, techincal leader and subject matter experts at Snowflake Summit to see how real-time feature pipelines power smarter ML. Schedule a Meeting with Chalk SCHEDULE A MEETING terminal Technical meetings Deep-dive sessions with Chalk engineers — architecture reviews, integration workshops, proof-of-concept discussions. user Executive meetings Strategic conversations with Chalk leadership — roadmap, partnership, and go-to-market alignment. The best Summit conversations happen after hours. Join Chalk at Joyride Pizza for an opening night kickoff party, located right outside Moscone North. Good pizza, better conversations, and data and AI builders worth meeting. Spots are limited — RSVP now. Joyride Pizza - Yerba Buena Gardens JUNE 1, 2026 / 6:00-9:00 PM Chalkin’ it out @ Snowflake Summit RSVP for the Chalk Kickoff Party! RSVP NOW kickoff party Chalkin’ It Out @ Snowflake Summit Join Chalk and friends at the biggest partner event of the week: a 1000+ person happy hour complete with live music, food trucks, beer gardens, games, and networking. Spark Social June 2, 2026 / 6:00-10:00 PM Sigma Social: Partners in the Park SPONSORED EVENT Experiences Some of the best moments at Summit happen off-schedule. Chalk's after hours events are where data and AI teams meet, swap war stories, and maybe grab a drink. No big productions — just the right people in the right room. Learn more about who we are, what we do and why it matters for teams operationalizing AI at real-time scale. Chalk 101 See how Whatnot uses Chalk to power real-time recommendations across its marketplace. Customer story Chalk was recognized as one of the world’s most innovative companies of 2026 by Fast Company. World's Most Innovative Discover how Chalk helps teams build real-time models for mission critical operations. See Chalk in action Explore more of Chalk Meet the team Marc spent many years at Google where he helped to launch the first version of Google Wallet. He went on to start Index, which Stripe acquired as its in-store payment solution—now called Stripe Terminal. Elliot started his career at Affirm where he built the early risk and credit data infrastructure system (the inspiration for Chalk). He then co-founded Haven Money, which Credit Karma acquired to power its banking products. Andy worked at Palantir on large government data infrastructure projects. He then co-founded Haven Money (with Elliot), which now powers Credit Karma Money. Melanie manages the post-sales forward deployed engineering and customer success teams at Chalk. Melanie previously worked on data infrastructure at Plaid, in addition to time at Airbnb, Two Sigma, and several founding teams. Alex is a seasoned enterprise SaaS GTM leader, scaling marketing and revenue teams from Series A through C across categories including AI, HRTech, and security. Her expertise has helped build beloved brands and turn pipeline engines into category-defining growth machines. Kyle is a GTM leader who has built post-sale organizations at category-defining companies — including Sprinklr and Databricks — and now applies that experience to helping customers successfully deploy real-time AI infrastructure. Caprice has spent 10+ years partnering with companies of all sizes helping them unlock lasting value from the technology they invest in. She brings a people-first approach to account management and a has a track record of building long-term relationships. Chuck has been leading recruiting teams at early stage startups for 10+ years with a a focus on high talent bars and scalability, notably at Plaid, AngelList, Block (formerly Square), and StubHub. Our booth number is #2611, located on the Moscone South side. Where is the booth? Please complete the form, indicate which type of meeting you're looking for and our team will reach out within 24 hours to schedule. How do I schedule a meeting? Yes, space will be limited, so we recommend registering now! Is the kickoff party open to everyone? For on-site demos, you can either request a meeting or swing by our booth at your leisure. For demos outside of Summit, please visit our website to schedule time with our technical experts. How do I get a demo? FAQ # The Platform for Production AI source: https://chalk.ai/landing/ship-production-AI-faster Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to deploy production-grade AI? Book a demo and see how Chalk powers real-time ML. # More than a feature store, the data platform for AI + ML source: https://chalk.ai/landing/ai-data-platform Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to Ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # More than a feature store, the data platform for AI + ML source: https://chalk.ai/landing/fintech-devcon Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to Ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Serve fresh, complex features sub 10ms, without streaming complexity source: https://chalk.ai/landing/serve-features-sub-10-ms Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to Ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # The ML stack is broken where it matters most: production source: https://chalk.ai/landing/ml-stack-is-broken Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Your models deserve better than stale features Talk to an engineer and see how Chalk can power your production AI and ML systems. # More than a feature store, the data platform for AI + ML source: https://chalk.ai/landing/snowflake-summit Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to Ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # More than a feature store, the data platform for AI + ML source: https://chalk.ai/landing/ai-council-one-page # More than a feature store, the data platform for AI + ML source: https://chalk.ai/landing/data-platform-for-ai-ml Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Trusted by leading ML teams for real-time feature engineering source: https://chalk.ai/landing/trusted-by-ml-teams Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Production ML should be this fast Talk to an engineer and see how Chalk can power your production AI and ML systems. # AI and ML data, on demand. source: https://chalk.ai/landing/ai-ml-data-on-demand Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Ready to Ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Real-time features for GenAI workflows source: https://chalk.ai/landing/real-time-gen-ai Used by ML and data platform teams at Chalk delivers the platform modern ML teams actually need One system for building, computing, and serving production-grade features Connect to any data source The latest at Chalk Make your GenAI application production-ready Talk to an engineer and see how Chalk can power your production AI and ML systems. # Ship Better Models: A New ML Workflow with a Context Layer and Agents source: https://chalk.ai/watch/use-chalk-notebooks-for-agents Join us for a session on rebuilding the ML workflow around a real-time context layer and an AI agent that helps data teams investigate, iterate, and ship better models. On-demand Webinar In this session, we’ll show how that same layer becomes the foundation for AI agents that can investigate model behavior, uncover missed opportunities, and run controlled experiments in a Chalk Notebook - using the trusted context your ML teams already rely on. We'll demonstrate how to: - Define and organize the features and real-time signals used by ML models - Initiate an ML data agent that retrieves and applies shared feature definitions - Continuously improve ML models through dataset generation, training, and experimentation in a Chalk notebook - Reduce duplicated logic across data science, ML engineering, and agent teams Define the features and real-time signals your models run on, then give an AI agent that same context so it can explain results, investigate issues, and find ways to improve ML performance. On-Demand Webinar # Life After Tecton Migrating to Chalk source: https://chalk.ai/watch/databricks-tecton Check out this on-demand webinar learn about your options following Tecton's acquisition by Databricks. Tecton’s acquisition by Databricks has left many ML teams uncertain about the future of their feature platform. Customers are asking what this means for the products they rely on, the complexity of migrating, and potential lock-in to the Databricks ecosystem. Check out this on-demand session to hear how leading ML teams are preparing for what’s next, and why many are moving to Chalk, a modern alternative built for control, speed, and interoperability. What you’ll learn: - Facing uncertainty: Key questions and early concerns following the acquisition. - A better path forward: Chalk runs in your cloud, connects to your existing ecosystem (Snowflake, Databricks, BigQuery, Postgres, Kafka, and more), and keeps your data, compute, and serving choices open. - White-glove migration: How Chalk’s engineering team ensures a seamless transition. Leading ML and data teams are already taking action. Watch now to explore your options and take back control. What Databricks’ acquisition means for your ML stack — and how to take back control. On-Demand Webinar Thanks for your interest! Watch the on-demand webinar below. # How Chalk fuels Whatnot’s live shopping platform source: https://chalk.ai/watch/whatnot-case-study Chalk powers Whatnot’s real-time recommendation engine, enabling low-latency personalization and scalable ML infrastructure for the largest live shopping marketplace. Whatnot is the largest live shopping platform in the U.S. and Europe, a marketplace where buyers and sellers connect through real-time, community-driven commerce. The company surpassed $3 billion in live sales last year alone, with users spending on average 80 minutes a day browsing and buying from live shows. The Data & AI team began rebuilding the recommendation system to support real-time online inference. They needed infrastructure that could serve tens of thousands of features per request at low latency, across deeply nested graphs and high-throughput workloads. They chose Chalk to power all core recommendation systems, from feed ranking to show discovery. # Chalk for Startups source: https://chalk.ai/startups Build and scale AI/ML products with Chalk for Startups: real-time data infrastructure, hands-on engineering support, preferred pricing, and GTM amplification for early-stage teams. Apply Access to hands-on forward-deployed engineers, preferred pricing, GTM amplification and a founder community so you can move faster and scale with Chalk: the platform to build and ship with real-time data. Backed by leading investors Work directly with Chalk's forward-deployed engineers, the people closest to the platform. Real answers and real help getting your models and agents into production. Hands-on forward-deployed support Bring us your architecture. We'll review your approach to real-time data and help you design for scale. Bring us your architecture. We'll help you design for scale, with preferred pricing built to support you from first deployment through hyper-growth. Preferred pricing When you launch or raise, we help you reach a wider audience of AI builders and founders. GTM amplification Join a community of technical founders solving hard problems, with events and introductions that compound over time. Founder network and events What accepted teams get More than infrastructure. You get the platform and resources to build faster, iterate in production, and spend less time on the work behind the scenes. Chalk for Startups is for technical teams building the next generation of applications. Apply if: - AI or ML is core to what you're building. - You've raised funding - Seed, Series A or B. - You're new to Chalk. Not sure you're a fit? Start a conversation and we'll figure it out together. Who can apply How it works Tell us about your team and what you're building. It takes a few minutes. Apply. We talk through your stack, your real-time use cases, and where we can help. Scoping Call. Accepted teams get up and running fast. Put your first real-time workloads into production on Chalk, and start building with our engineers alongside you. Your first 30 days. Chalk for Startups is designed to meet you at the build moment and scale with your team. We'll walk through the details on your scoping call. Built to grow with you Explore all of Chalk's solutions Technical teams building AI or ML into the core of their product, from your first raise through Series B. If real-time data is central to what you're building, we want to talk. Who is Chalk for Startups for? Direct access to Chalk's forward-deployed engineers, a preferred pricing model, design review, GTM amplification on your launches and raises, and a founder network. What do accepted teams get? No. This is a hands-on program. You apply, we talk through fit, and accepted teams work directly with our team from day one. Is this self-serve? Chalk helps your models and agents act on fresh data straight from data warehouses, APIs and streams by providing a federated query engine and infrastructure platform that runs in your own cloud. Your data and infrastructure stay under your control. What does Chalk do? A few minutes. Tell us about your team and what you're building, and we'll follow up to set up a scoping call. How long does the application take? No. The program is for teams new to Chalk. Do I need to already be a Chalk customer? FAQ Ready to apply? Tell us what you're working on and we'll take it from there. # Chalk for AI Platform Engineers source: https://chalk.ai/for/platform-engineers Chalk gives platform engineers one end-to-end system to build, test, evaluate, deploy, observe, and improve AI systems. TALK TO AN ENGINEER Chalk for Platform Engineers Power the full development cycle for production AI with the Chalk platform. Trusted by teams building the next generation of AI / ML Why AI platform engineers choose Chalk Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project The workflow for building, testing, and deploying AI is fragmented. - Source data lives in one system. - Observability tools are in another. - Evaluation systems are somewhere else. - Production infrastructure, sandboxes, tool gateways are across vendors. - Everything has to be wired together. Chalk gives AI platform engineers one end-to-end system for building, evaluating, and running real-time, production AI. One continuous loop. Develop, simulate, evaluate, deploy, observe, and improve in one platform — without rebuilding the stack at every stage. Develop, simulate, evaluate, deploy, observe, and improve in one platform - without rebuilding the stack at every stage. One continuous loop. One execution path. Build, test, and run AI against the same underlying data, logic, tools, and infrastructure. Build, test, and run AI against the same underlying data, logic, tools, and infrastructure. One execution path. One system of context. Give agents the live business state they need without creating another copy of your data. Give agents the live business state they need without creating another copy of your data. One federated context layer. ALL ON ONE PLATFORM Chalk brings the workflow together. - AI Agents: Build agents with real-time context. Simulate, evaluate, and deploy within one platform. - LLMs: Train and fine-tune open-weight models, evaluate performance against a specific task, and then serve in your cloud. - Real-time ML models: Define features, query across sources, compute and serve decision-ready inputs in milliseconds, and run your model where your data lives. Built for the continuous loop of AI development. Ready to ship next‑gen AI? Talk to an engineer and see how Chalk can power your production AI systems. # Chalk for MLOps source: https://chalk.ai/for/mlops Make ML production-grade with Chalk: reproducible, governed, and reliable across batch and real-time. TALK TO AN ENGINEER Build fully observable ML pipelines with automatic data lineage on every query. Trusted by teams building the next generation of AI + ML End-to-end visibility into every feature’s query plan, upstream inputs, and execution path Full visibility into how every feature is built for better debugging Data lineage and traceability Automatic versioning and audit trails for consistent, reproducible feature definitions Built-in versioning and audibility Feature observability, monitoring, and alerting for freshness, drift, and quality Monitoring with alerts so features remain fresh, accurate, and drift-free Feature monitoring and alerting Logs for debugging, Kubernetes cluster activity, and system-level metrics with alerts Infrastructure observability Automatic orchestration that replaces fragile pipelines On-demand pre-computation that replaces brittle pipeline logic and manual orchestration On-demand pre-computation Sub-5ms real-time serving on aggregations at high throughput Sub 5ms real-time serving at high throughput Low-latency serving at scale Why MLOps engineers choose Chalk Tracing gives teams deep visibility into how queries run inside Chalk. Each resolver and model call is instrumented and timed, making it easy to identify performance bottlenecks and understand why a query behaves the way it does. Tracing docs Tracing for query performance diagnosis Chalk makes it easy to get started, and its isolation model keeps everything safe by default. Teams don’t block each other anymore. Chalk applies modern software engineering to ML workflows: - View lineage from data sources through transformations to final outputs, with version history captured automatically - Browse and search historical versions of every feature, query, resolver, and deployment in one unified feature catalog - Audit changes and reproduce results instantly with complete traceability built into the dashboard MLOPS IN ACTION Reproducible and auditable features by default Chalk runs in your cloud (AWS, GCP, Azure), meeting enterprise standards for security, compliance, and deployment flexibility. DEPLOY IN YOUR CLOUD Built for scale and trust Explore how Chalk works Ready to ship next-gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Chalk for Data Scientists source: https://chalk.ai/for/data-scientists Empower data scientists to build reproducible, production-ready ML features. Chalk unifies training and inference for faster experimentation. TALK TO AN ENGINEER Discover and create features, generate reproducible datasets, and accelerate AI and ML experiments with Chalk. Trusted by teams building the next generation of AI + ML Data Scientists move quickly from hypothesis to production. With Chalk they: - Get observability into the features they serve - Build training datasets with lineage - Trust that training and production always stay in sync Learn More Why data scientists choose Chalk By moving our feature pipelines to Chalk, Data Science and Engineering now work side-by-side throughout model development. What used to be lengthy, error-prone handoffs are gone. Our entire search ranking stack runs on Chalk, serving features for inference in under 50ms, and we’re extending it to all of our models, including real‑time personalization. Training data should reflect what your application would have known at the time. Chalk guarantees this with point-in-time lookups: - Avoid leakage by using only values available at the correct historical moment - Generate reproducible datasets with complete lineage - Backtest models with confidence, knowing inputs match production LEARN MORE Point-in-time correctness @features class Review: id: int at: FeatureTime product_id: "Product.id" review_body: str rating: int @features class Product: id: int title: str reviews: DataFrame[Review] average_rating: float | None = _.reviews[_.rating].mean() last_rating: int | None = _.reviews[_.rating].max_by(_.at) review_count: Windowed[int] = windowed( "30d", "90d", "all", expression=_.reviews[ _.created_at > _.chalk_window ].count(), ) python Chalk integrates directly with the tools data scientists use every day, enabling them to explore, iterate, and ship models without leaving their notebook. - Iterate on features and model training directly in your notebook - Deploy, backfill, and monitor datasets for training pipelines - Experiment and define new features independently while Chalk handles orchestration Notebook tutorial Built for notebooks Explore how Chalk works Ready to ship next‑gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Chalk for AI + ML Engineers source: https://chalk.ai/for/ml-engineers Power models with production-grade infrastructure, all from a single Python interface. Experiment faster, unify structured and unstructured data, serve features in real time with Chalk. Experiment faster, unify structured and unstructured data, serve features in real time, and power models with production-grade infrastructure, all from a single Python interface. Trusted by teams building the next generation of AI + ML Work with structured data like transactions and aggregations, and unstructured inputs like text, embeddings, and prompts Unified data for every modality Experiment faster by defining features once in Python and reusing them across training and inference Faster experimentation loop Generate reliable training data with point-in-time correct backfills Reliable, point-in-time datasets Serve models in real time with sub-5ms feature retrieval from heterogeneous sources Serve models in real time with sub 5ms feature retrieval from heterogeneous sources Low-latency model serving Power LLMs with embeddings, vector search, and prompt construction LLM-native feature workflows Reproduce and monitor all queries with built-in versioning and query logs Reproducible by default Why AI/ML engineers choose Chalk Chalk powers our LLM pipeline by turning complex inputs like HTML, URLs, and screenshots into structured, auditable features. We can serve lightweight heuristics up front and rich LLM reasoning deeper in the stack, catching threats others miss without compromising speed or precision. @features class Product: id: str description: str embedding: Vector = embed( input=lambda: Product.description, provider="vertexai", model="text-embedding-005", ) @features class User: id: str purchases: DataFrame[Purchase] embedding: Vector = _.reviews[_.product.embedding].mean() recs: DataFrame[Product] = has_many( lambda: Product.embedding.is_near( User.embedding ) ) python LLM toolchain extends feature engineering into the world of GenAI - Embeddings support: define, store, and retrieve embeddings as features - Prompt construction: build structured, dynamic prompts using Chalk features - Vector search: retrieve semantically relevant context at query time With Chalk, AI/ML engineers can combine structured and unstructured data for more accurate, context-rich LLM applications. EXPLORE LLM TOOLCHAIN LLM-ready infrastructure Chalk’s compute engine, powered by Velox, delivers vectorized, low-latency execution across both batch and real time. Features are served in less than 5ms, even with complex transformations and joins. REAL-TIME SERVING Built for performance + scale @features class Website: id: int content: str url: str completion: P.PromptResponse = P.completion( model="gpt-5.1-2025-11-13", messages=[P.message( role="user", content=F.jinja(""" Analyze the following website: url: {{Website.url}} content: {{Website.content}}""")) ], # Use structured output as dataclasses output_structure=CompanyCompletion, ) Chalk unifies structured and unstructured data from the start of model development. This single definition can be: - Backfilled into training datasets - Served in real time at inference with millisecond latency - Versioned and reproduced at any point in time Chalk ensures models are trained and deployed on the same feature logic, eliminating drift. From training to real‑time inference Explore how Chalk works Talk to an Engineer Ready to ship next‑gen AI/ML? Talk to an engineer and see how Chalk helps AI/ML engineers deliver faster experimentation, real‑time inference, and LLM‑powered applications. # Chalk for Data Engineers source: https://chalk.ai/for/data-engineers Chalk helps data engineers define, compute, and serve features in Python—ensuring consistency across batch, streaming, and real-time ML pipelines. TALK TO AN ENGINEER Chalk unifies your data schema to compute and serve features defined in Python consistently across batch, training, and real time. Trusted by teams building the next generation of AI + ML Chalk automatically builds and executes DAGs for computing your queries, removing custom orchestration Eliminate manual pipelines Ensure consistent feature definitions across backfills, training sets, and real-time inference from a single source of truth Unified feature catalog Define rolling windows, decays, and normalized features directly in Python Simplify temporal aggregations Low-latency APIs without Kafka, Flink, or custom streaming stacks Serve in real time without extra systems Create point-in-time-correct datasets with full lineage and versioning Generate reproducible training data Why data engineers choose Chalk @online def get_quote_is_risky( owned_vehicles_count: Quote.owner.owned_vehicles_count, n_addresses_30d: Quote.owner.n_addresses["30d"], n_addresses_1yr: Quote.owner.n_addresses["365d"], ) -> Quote.is_risky: return ( n_addresses_30d > 1 or n_addresses_1yr > 5 ) and owner_owned_vehicles_count > 2 Chalk isn’t just a feature store. It’s an execution graph for your features. At request time, Chalk slices the graph and computes only what’s needed, from the freshest data available. EXPLORE ONLINE QUERIES Define once, use everywhere Instead of relying on stored data to stay in sync, we compute directly from the source. Chalk makes that both precise and reliable. Explore how Chalk works Ready to ship next‑gen ML? Talk to an engineer and see how Chalk can power your production AI and ML systems. # Build agents with context to solve support tickets quickly. source: https://chalk.ai/use-case/customer-support-agents Build customer support agents on live account, transaction, and interaction data. Chalk serves fresh context at inference with governed access and millisecond latency. TALK TO AN ENGINEER CUSTOMER SUPPORT AGENTS Chalk serves fresh account, transaction, and interaction data to customer support agents at decision-time. Agents handling sensitive customer issues get the current information they need, so resolution is smooth, fast, and accurate for your customers. TRUSTED BY For ML + AI engineers Test agents against what actually happened before putting them in front of customers. Chalk replays historical tickets and interactions in sandboxes with an enforced knowledge cutoff. The agent sees only the information that would have been available at that point in time. Ship support agents you can trust. Support conversations change by the second. Chalk resolves data across streams, databases, warehouses, APIs and other sources when the agent needs it, so every decision can use current information. Give agents live context for accurate ticket resolution. Restrict agent context to the data it needs to solve the issue. Enforce Rego policies on every MCP tool call, making decisions based on the user, scopes, backend, tool, and tool-call arguments. The agent gets the access it needs - and nothing else. Keep the agent’s access to customer data in scope. For product + platform leaders Reduce time spent on simple case work and use agents to cover more ticket volume, so your team can focus on high-judgement cases. Solve more support cases without scaling headcount. Voice and chat interactions complete in seconds end-to-end. Chalk delivers low-latency context, so customers get the help they need quickly and fewer sessions are abandoned. Provide a seamless, fast support experience. One customer's data never reaches another customer's chat. Register approved tools with the MCP Gateway, authenticate each sandbox through its workload identity, and route tool calls through a governed proxy for centralized policy enforcement and auditability. Ensure security at the infrastructure layer and avoid a PR disaster. Tailor agents for every stakeholder Customer-facing resolution agents Define the features and data your agents need in Chalk, then serve governed context at inference-time. Every query can be traced back to the data used to produce an answer. "My order never arrived. Can you give me a refund? " Contact center agents Compute live intent and account features from the interaction in progress and serve them, fast. "I’m writing in because I want to know the status of my shipment. " Internal support triage Give investigation agents governed access to the data and tools needed to work alongside your team. The agent does the investigation. Your review team makes the final call. "Seller says the campaign underperformed. What happened? " Additional resources BLOG /blog/chalk-compute Introducing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud. https://chalk.ai/workflow/evaluate-agents Discover how to evaluate and improve agents on Chalk https://docs.chalk.ai/docs/compute/sandbox Learn More about Chalk Compute sandboxes https://docs.chalk.ai/docs/security Learn more about Chalk’s security architecture Bring us the customer support agent that hasn’t learned your company’s workflows. We’ll show you how to build, evaluate, and monitor governed, context-aware agents, directly in your cloud. # Build agents that operate with live context and help your team get work done. source: https://chalk.ai/use-case/workflow-automation-agent Build internal workflow agents that act on live business data with governed access and audit trails. Run agents in isolated sandboxes with fresh context and millisecond-latency queries. TALK TO AN ENGINEER WORKFLOW AUTOMATION AGENTS Process work like invoice review and status reports is repetitive. Agents can take on these workflows when they can safely access fresh business data. Chalk helps you build, test, and deploy agents in sandboxes inside your VPC, with governed access to live data from the systems where it lives. For ML + platform practitioners Internal agents stall on stale data. Chalk federates queries directly across your data warehouse, data lake, streams and APIs at inference time. No need to set up ETL jobs. Give agents live data. Grant access to specific tables, columns, and rows. Enforce Rego policies on every MCP tool call, making decisions based on the user, scopes, backend, tool, and tool-call arguments. Define policies so you know what the agent can and can’t touch. Control exactly what agents can access. Agents that work with your proprietary and production data should stay in your environment. Chalk runs them in isolated sandboxes inside your VPC. Run agents securely in your cloud. For engineering + platform leaders Ground every agent action in current data. Simulate a historical point in time and evaluate agents before rolling them out. Power your team with agents they can trust. Reduce repetitive investigation work like searching reports, querying records, and tracking issues with agentic workflows. Put engineering time where it matters. Agents assemble the case file, run standard checks, resolve routine cases, or hand a reviewer a decision-ready summary. Automate repetitive business processes. Finance Invoice processing, expense review, and revenue reporting are repetitive - and require access to sensitive business data. Build, test, and deploy agents in sandboxes inside your VPC, with governed access to live data across your systems, and a complete audit trail. Make financial analysis an agentic workflow. Operations Reporting Checking vendor history, validating purchase orders, and flagging anomalies are manual work - and require access to sensitive business data. Build, test, and deploy agents in sandboxes inside your VPC, with governed access to live data across your systems, and a complete audit trail. Turn recurring reporting into an agentic workflow. Product & Growth Experiment analysis, cohort readouts, and feature adoption summaries aggregate data into insights. Build, test, and deploy agents in sandboxes inside your VPC, with governed access to live data across your systems, and a complete audit trail. Get experiment insights with an agentic workflow. Additional Resources https://docs.chalk.ai/docs/compute/sandbox Learn More about Chalk Compute sandboxes How a single knowledge cutoff locks an entire agent trajectory to a point in time. Read the announcement → BLOG /blog/chalk-compute Introducing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud. Docs https://docs.chalk.ai/docs/compute/overview Learn about Chalk Compute Bring us the processes that slow your team down. We'll show you how to simulate, evaluate, and deploy your agent with controlled access to the data that makes it useful - all in your cloud. # Trust every identity decision with real‑time fraud signals source: https://chalk.ai/use-case/identity-fraud Power real-time identity fraud detection with fresh, point-in-time correct ML features for risk scoring and abuse prevention. TALK TO AN ENGINEER IDENTITY FRAUD Chalk computes and serves identity, device, behavioral, and KYC features in real time. Risk models can verify users, prevent account abuse, and stop synthetic identities in single-digit milliseconds. Trusted by risk and identity leaders worldwide ML + Data Practitioners Define velocity checks, device fingerprints, behavioral features, and document validity directly in Python. Ship new risk signals fast Combine device, behavioral, onboarding, KYC, and third-party signals into a single, consistent feature set. Unify identity data Inspect every feature and execution path with Chalk’s query plan viewer. Debug high-stake decisions FRAUD + RISK LEADERS Eliminate latency and stale features. Power models with fresh signals like user actions, API responses at inference time. Power real-time identity decisions Use live identity context for better decisions while minimizing unnecessary step-ups and user drop-off. Reduce risk without slowing growth Monitor outputs with full observability; built-in lineage, audit logs, and PII masking. Seamless compliance Why Chalk for fraud and identity Behavioral abuse and account risk Verisoul uses Chalk’s Python-native features and low-latency serving to update fraud detection logic 10× faster and stop attacks as they happen. Detect anomalous user behavior and activity velocity in real time. Account Takeover Customers use Chalk's feature platform to access real-time device and behavioral signals for blocking account abuse, served with single-digit millisecond latency. Stop bots, fake accounts, and synthetic logins. Onboarding & KYC Mission Lane cut feature deployment cycles from weeks to days by using Chalk to unify credit bureau and onboarding signals into consistent features. Verify documents and identities with fresh, unified data. Synthetic IDs & emerging fraud MoneyLion uses Chalk’s versioned features and backtesting to validate new fraud signals on historical data and ship updated logic to production in hours. Detect synthetic identities and adapt to emerging attack vectors from chalk import online @online def get_username(email: User.email) -> User.username: username = email.split("@")[0] if "gmail.com" in email: username = username.split("+")[0].replace(".", "") return username.lower() @features class Transaction: id: int amount: float user_id: "User.id" python Fetch device info, IP geolocation, velocity checks, and KYC scores at inference time. No batch lag, no stale features. Fraud Detection Tutorial Real-time risk features in single-digit millisecond latency Verisoul uses Chalk to detect bots and coordinated fake accounts in real time. With Chalk they ship fraud features 10× faster, improve detection accuracy 4× with inference-time data, and maintain full auditability across their models. READ STORY Stopping faking accounts with fresh fraud features Additional resources Build identity fraud systems powered by real-time data Get started with Chalk today and build risk infrastructure that moves at the speed of your business. Explore more of Chalk’s solutions # Power search and ranking with real-time features source: https://chalk.ai/use-case/search-and-ranking Build search and ranking models with real-time features, embeddings, and vector search. Re-rank results instantly with live context using Chalk. TALK TO AN ENGINEER SEARCH AND RANKING Chalk powers search models by computing relevance features and embeddings at query time. Teams leverage Chalk to adapt results instantly based on user context, behavior, and intent, while maintaining full control over ranking logic and data. Trusted by leading teams building large-scale ranking systems ML + data practitioners Retrieve candidates using embedding similarity, and re-rank results with live behavioral and contextual features at query time. Retrieve and rank with embeddings in real time Compute features and embeddings on demand with predictable, single-digit millisecond latency, supporting high query volume and large candidate sets. Serve search features with ultra-low latency Use the same feature definitions for training and production search traffic to prevent online and offline drift as ranking logic evolves. Maintain point-in-time correctness Search and product leaders Deliver more relevant results by ranking with embeddings and live context computed at query time, not static scores generated hours earlier. Improve relevance with decision-time ranking Change ranking logic and features without rebuilding batch jobs or maintaining separate online and offline paths. Adapt rankings without pipeline rewrites Serve billions of searches with consistent latency while validating ranking changes before they reach production. Scale search without relevance regressions Why Chalk for search and ranking models Chalk’s performance directly affects the quality of our search and discovery models, which power everything from price flexibility to apartment ranking. The ability to call real-time features without dealing with stream complexity has been huge for us. ecommerce and content discovery Retrieve candidates using query and item embeddings, then re-rank results with live behavioral and contextual features so relevance reflects meaning, not just keywords Rank results based on semantic relevance. Marketplace search Compute relevance, availability, and demand features at query time to rank listings based on what is actually available right now, not static indexes or delayed aggregates. Rank results using live supply and availability. Geospatial search Rank listings using real-time user preferences, location context, and session behavior so results adapt instantly as users refine filters and searches. Personalize GIS results using live user intent and geospatial data. Whatnot delivers dynamic recommendations during live streams with Chalk, serving user and product features instantly to maximize engagement. Live recommendations at scale whatnot preview Run nearest-neighbor search on embeddings at query time and re-rank results using the same production feature definitions. Nearest neighbor docs Vector search with Chalk Apartment List uses Chalk to compute ranking features at query time, keeping search results aligned with live user intent and inventory changes. This allows the team to personalize search results, iterate on ranking logic safely, and maintain consistency between offline evaluation and production behavior. Read story Search and ranking at scale Additional resources Build search systems that adapt instantly. Get started with Chalk today and transform your ML workflows. Explore additional use cases # Price what is happening now. source: https://chalk.ai/use-case/dynamic-pricing Chalk serves pricing models fresh demand, supply, and behavioral features at inference-time. TALK TO AN ENGINEER DYNAMIC PRICING Chalk serves pricing models fresh demand, supply, and behavioral features at inference-time. Increase revenue and optimize demand with real signals at the moment price is decided. ML + Data Practitioners Chalk resolves features with data from streams, databases, and APIs at query time, serving them in single-digit milliseconds. Reduce your pipeline complexity Define features in Python or SQL, and deploy on your own cadence. Turo cut new feature delivery from 3 weeks to 1 week. Ship new features faster Compute features when the model runs - and make real-time signals available for pricing and the decisions that depend on what is happening now. Serve fresher data at inference For product + platform leaders The first model Medely deployed on Chalk, a charge-rate recommender, added $800K in annual net revenue. Build models that drive revenue Iterate with feedback models from development to production in days, not quarters. Apartment List cut model deployment from weeks to 1-2 days, creating more time for pricing experiments. Improve model performance Shorten the path from hypothesis to production. Turo cut feature development and experiment cycles by 67%, giving the team more opportunities to test and improve pricing. Run experiments faster Why Chalk Marketplace Pricing Turo computes host pricing recommendations from live marketplace conditions across 300,000+ vehicles in five countries - and ships new pricing features in a week instead of three. Supply and demand shift by neighborhood and by hour. Rate recommendations Medely recommends charge rates for healthcare shifts using current demand, fill rates, and market conditions. Features that arrived on a 24-hour batch lag now resolve in real time at inference. The result: $800K in projected annual net revenue from the first model. Features resolve in real time at inference Price flexibility and ranking Apartment List powers price flexibility and apartment ranking models with sub-10ms dynamic recommendation queries. The models deliver personalized results while the renter is still searching, on features retrieved in under 5ms. Models deliver personalized results from chalk import online @online def get_username(email: User.email) -> User.username: username = email.split("@")[0] if "gmail.com" in email: username = username.split("+")[0].replace(".", "") return username.lower() @features class Transaction: id: int amount: float user_id: "User.id" python Fetch device info, IP geolocation, velocity checks, and KYC scores at inference time. No batch lag, no stale features. Read the Docs Real-time risk features in single-digit millisecond latency How Medely's first model on Chalk added $800K in projected annual net revenue The first product we ever deployed with Chalk paid for our team, probably more, in net revenue. Adopting Chalk is the biggest singular win I have had as an ML engineer at this company. Medely matches healthcare professionals with open shifts, and the charge rate on every shift is a pricing decision. Before Chalk, the features behind those decisions arrived on a 24-hour batch lag. Now they resolve in real time at inference, and experiment cycles have dropped from roughly two months toward two weeks. READ STORY Bring us your pricing decisions that need fresh data Every pricing team has a model waiting on batch data. Bring us yours. We'll show it scoring on live demand, supply, and behavior, running in your stack. # Build agents that manage abuse detection with real-time context. source: https://chalk.ai/use-case/trust-and-safety-agents Build agents that detect abuse, investigate threats, and take action using real-time user and platform context. Sandboxed execution inside your cloud. TALK TO AN ENGINEER TRUST & SAFETY AGENTS Abuse investigation, infringement detection, and cyber defense depend on fast access to the right context. Chalk gives agents real-time information like behavioral signals, user activity, messages and other data from across your stack, plus a secure sandbox to investigate and act. Trusted by leading companies For AI + platform engineers Ship new detection signals faster and improve decisions with inference-time context from across your data stack. Verisoul ships fraud features 10x faster on Chalk and improved detection accuracy 4x. Ship new detection signals faster and improve decisions with inference-time context from across your data stack. Verisoul ships fraud features 10x faster on Chalk and improved detection accuracy 4x Add signals without rebuilding infrastructure. Record the context an agent used, evidence it found, tools it called, and action it recommended or took. Review outcomes, improve policies, and maintain accountability. Make every agent decision auditable. Chalk runs agents in isolated sandboxes inside your VPC, with controlled network access and permissions scoped to the case. Run agents safely. For trust and safety leaders Agents gather evidence, cross-check signals, apply policy, and automate repetitive investigation work. Cases move from queue to action without the day-long wait. Resolve routine cases faster. Connect activity across accounts, devices, identities, content and time. Surface emerging attack patterns before they scale. Detect coordinated threats earlier. Automate well-defined investigations and escalations so reviewers can spend their judgment on ambiguous, severe, and novel cases, with lower cost per decision. Focus your team’s judgement on the hardest cases. Targeted capabilities for your team Fake accounts and coordinated abuse Verisoul detects bots and coordinated fake account rings in real time on Chalk, shipping fraud features 10x faster and improving detection accuracy 4x with inference-time data. Threat investigation Agents gather account history, linked accounts, device overlap, payment records, and prior reports into a single view, then reason over it with tools scoped to what the case needs. CONTENT & BEHAVIOR CLASSIFICATION Run embedding and classification models on your own infrastructure, close to the features that give them context. ENFORCEMENT & POLICY ITERATION Agents suspend, restrict, or escalate through permissioned tools. Set which actions run automatically and which require a reviewer. Evaluate a policy change against historical cases before it reaches live users. How Grindr runs trust and safety agents inside their own cloud "Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project." Grindr serves over 15 million users. The data those safety decisions depend on is exactly the data that cannot leave their environment. They run Chalk Compute directly in their own cloud to orchestrate trust-and-safety workloads, Agents read live user and platform context through the same engine that serves their production models. Additional Resources Announcement /blog/chalk-compute Introducing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud Engineering /blog/chalk-prompt-evaluation Which LLM Wins at Nolan Trivia? Chalk's Prompt Evaluation in Production Blog /blog/what-is-a-context-engine What Is a Context Engine Give trust and safety agents real-time context. Bring us your agent proof-of-concept and we’ll chart the path to production. # Power recommender systems with real-time features source: https://chalk.ai/use-case/recommender-systems Build real-time recommendation systems with fresh features, low latency inference, and no training-serving skew. Power dynamic personalization with Chalk. TALK TO AN ENGINEER RECOMMENDER SYSTEMS Chalk powers personalized feeds by computing fresh user, item, and marketplace features at request time. Teams use Chalk to serve recommendations that reflect real-time behavior and semantic similarity across dynamic inventories. Trusted by leading teams building real-time personalization ML + Data Practitioners Define embeddings, affinity scores, and session features, and compute them at request time. Compute recommendation features in single-digit milliseconds Use clicks, views, searches, and interactions as real-time inputs to feed ranking and personalization. Serve live behavioral features Reuse the same feature definitions between training and inference to avoid inconsistencies in feeds. Keep training and serving consistent Recsys + growth leaders Serve millions of personalized experiences per second with predictable low latency as candidate sets grow and inventory changes rapidly. Scale personalization consistently Personalize content and products using live user context instead of static profiles. Deliver more relevant feeds Surface relevant items and sellers for new users and inventory using embeddings and request-time features. Handle cold starts in marketplaces Why Chalk for recommendation systems and personalization Whatnot delivers dynamic recommendations during live streams with Chalk, serving user and product features instantly to maximize engagement. Live auction recommendations at scale whatnot preview ecommerce Compute user, item, and session features at request time to rank products based on live behavior and semantic similarity. Personalize product and content feeds in real time. Financial services Compute fresh user and account features at request time to personalize product feeds and recommendations across changing financial context. Personalize offers and experiences using live account signals. Marketplaces Compute relevance and availability features at query time so listings reflect live inventory, location context, and evolving user intent. Adapt discovery as inventory and preferences change. Healthcare Use embeddings and request-time features to surface relevant content or care pathways based on recent activity and contextual signals, while reusing the same feature definitions for evaluation and production. Personalize recommendations using live context and embeddings. Match users to items or sellers using semantic similarity. Chalk computes embeddings and retrieves candidates via vector search, then personalizes results using real-time context, supporting large candidate sets without relying on static precomputed recommendations. Embeddings docs Embedding-based discovery Marketplaces introduce additional complexity. Inventory turns over rapidly. Buyers, sellers, and items constantly cold start. Signals are spread across multiple entities. Chalk models these relationships explicitly and computes features across them at request time, allowing personalization to adapt instantly as users interact, inventory changes, or new items are introduced. Query Planner Built for dynamic marketplaces Designing a feature platform that fits all team use cases, and still performs, was one of our biggest challenges. More Resources Build feeds that respond in real time Use Chalk to compute recommendation features when relevance matters most. Explore additional solutions # Build ML research agents with production context. source: https://chalk.ai/use-case/ml-autoresearch-agent Chalk gives agents the same context, production data, and feature definitions your models were trained on. Agents can research architectures, analyze model decisions, build features, evaluate results, and retrain models, creating a recursive self-improvement loop for the ML lifecycle. TALK TO AN ENGINEER ML AUTORESEARCH AGENT Trusted by teams putting agents and models in production For ML + data practitioners Automatically explore hypotheses - suggest an idea for a feature and have an agent automatically engineer features, backtest with point-in-time correct training sets, and push PRs. Concurrently run many investigations on burst-scalable infrastructure. Parallelize exploratory investigations. Agents can build features, evaluate model changes, and prepare retraining runs using the same data and feature definitions that power production. Automate the path from diagnosis to a tested fix. Your team sets the direction, and agents run the experiments. Each loop proposes a feature, backtests it on point-in-time correct training sets, and keeps only what improves the metric. Less manual investigation required to improve your production models. For ML + platform leaders Turn production signals into model improvements without waiting for an engineer to start the investigation - helping your team iterate faster and improve outcomes at scale. Increased model performance, faster Ground every agent action in current feature data, lineage, and model artifacts. Get agentic recommendations based on the same production context your models use, and increase confidence in what agents find and propose. Power your team with agents you can trust Reduce time spent on repetitive investigation work such as searching logs, querying data, and tracing feature issues with agentic feature development. Put engineering time where it matters Why Chalk Agentic feature development for data engineers Give agents access to production data, feature definitions, lineage, and model context. Ask what features could improve a model and where the underlying signals live. The agent investigates the data, proposes features, and helps turn them into production-ready pipelines. Build and ship features faster Agentic signal identification for data scientists Ask the agent which signals drive model behavior and where new predictive signals may exist. It analyzes feature values, distributions, and relationships across production data to surface promising signals. Move from hypothesis to validated feature faster. Find the signals that improve predictions Agentic model debugging for ML engineers When model performance drops, ask the agent why. Debug production models with the full context behind each decision. Find what changed in production Models run on live features and real-time signals. Watch Chalk co-founder Elliot Marx build an agent with the same context. It can investigate model decisions, trace production issues, and identify changes that can improve model performance. Create an agentic loop from investigation to improvement grounded in the data your models use in production. Watch here Watch on demand: Give your ML team a research agent with production context Automate your model development workflows Bring us the decision your team spent last week investigating. We'll show you how to build an agentic ML improvement cycle. # Real-time feature computation for payments source: https://chalk.ai/use-case/payments Build real-time payment decisioning systems with Chalk. Compute transaction risk features at authorization time with millisecond latency. TALK TO AN ENGINEER PAYMENTS Chalk computes feature values at authorization time so models can evaluate risk using fresh, decision-time data. Trusted by industry leaders and developers worldwide AI/ML PRACTITIONERS Introduce new features quickly and validate them against historical outcomes before they impact live authorization decisions. Ship new authorization features fast Reuse the same feature definitions for offline training and online inference to eliminate online-offline skew. Eliminate online and offline skew Inspect execution plans, data lineage, and computed feature values for every request. Inspect features FRAUD + PAYMENT LEADERS Reduce latency and stale features. Power models with fresh signals at inference time. Power real-time payment decisions Serve fresh features to improve model accuracy while maintaining high approval rates and a smooth customer experience. Reduce fraud losses without increasing false declines Monitor outputs with full observability, built-in lineage, audit logs, tracing, and PII masking. Seamless compliance For teams running payment models Chalk helps us deliver financial products that are more responsive, more personalized, and more secure for millions of users. It’s a direct line from infrastructure to impact. TRANSACTION FRAUD Compute features at decision time to evaluate spend behavior and velocity signals. Detect anomalous spend and transaction patterns at authorization time. CARD AUTHORIZATION RISK Compute hundreds of fraud features per transaction in single digit milliseconds at p99 latency, powering real-time authorization decisions at scale. Score risk in the most latency-sensitive moment of the payment flow. ACCOUNT-LEVEL SPEND ABUSE Compute rolling and historical aggregations on demand to detect patterns. Identify coordinated abuse across transactions. EMERGING PAYMENT FRAUD Backtest new features and deploy the same feature definitions to production without rebuilding pipelines. Adapt quickly as fraud patterns change. Power real-time fraud systems where hundreds of feature values must be computed per decision under strict latency, reliability, and compliance constraints. Teams use Chalk to support authorization-time decisioning while maintaining consistency between offline validation and production behavior. Real-time serving Built and proven in global payment systems Chalk deploys within your cloud. All feature computation runs alongside your data sources, enabling data residency, compliance, and integration with your IAM and networking controls. Deploy in your cloud Enterprise-grade security By unifying offline and online computation, teams iterate faster, eliminate online/offline skew, and operate fraud systems that continuously improve without added infrastructure complexity. Ship better fraud decisions without risking approvals More resources Build payment fraud systems for the moment decisions matter Use Chalk to define, compute, and serve fraud features at authorization time. Move faster, reduce risk, and maintain control as fraud patterns evolve. Explore additional use cases # Underwriting demands real-time data source: https://chalk.ai/use-case/underwriting Build real-time credit underwriting models with accurate, auditable data. Chalk powers credit decisions with point-in-time correctness. TALK TO AN ENGINEER UNDERWRITING Underwriting decisions depend on fresh, high-conviction data. Chalk is the federated query execution engine behind modern credit and underwriting teams. It computes features at decision time, enables faster iteration, and scores applications in single-digit milliseconds at scale. Trusted by credit leaders ML and data practitioners Compute features like FICO blends, inquiry velocity, delinquency counts, and balance trends using decision-time data instead of batch aggregates Ship new credit signals fast Compute credit features across bureau, banking, and application data including cash-flow and account-level signals from providers like Plaid Unify credit data Use the same feature definitions for training, backtesting, and live scoring Deploy consistently Risk and decision leaders Decision with live credit bureau, banking, and and third-party signals instead of stale aggregates Improve approval accuracy Detect recent delinquencies, inquiries, cash-flow changes, and balance changes in real-time Reduce portfolio risk before approval Operate with built-in lineage, audit logs, and versioned feature definitions Meet compliance requirements Why Chalk for underwriting models Mission Lane leverages Chalk to unify offline and online features, cut deployment time from weeks to days, and move to real-time credit decisions. READ STORY Credit decisions at scale Buy Now, Pay Later Compute underwriting features using banking, cash-flow, and identity signals at decision time, enabling approvals that do not rely solely on traditional credit scores. Make instant credit decisions with alternative data Cash advance Evaluate recent income, account balances, and spending patterns using decision-time features, which is critical for cash advance products where FICO alone is insufficient. Assess short-term credit risk using real-time financial behavior. Business lending Compute features from bank accounts, revenue forecasting from Stripe, and Quickbook expense data to assess risk and approve loans for businesses. Underwrite businesses using a live view of financial health. Card issuing Combine bureau data with recent banking and transactional signals to determine eligibility and limits for new card products at decision time. Approve access and set limits at the moment of issuance. Version every feature, audit every change, and trace data lineage automatically. Chalk ensures every underwriting decision is transparent and regulator-ready. Lineage Docs Compliant credit decisions Fetch bureau data, banking signals from providers like Plaid, and alternative sources at decision time. Combine streaming and historical data to support accurate approvals in emerging credit products where traditional bureau data falls short. Real-time credit signals Every feature has to be precise. A single misaligned timestamp can change a lending decision. Chalk’s feature engine gives us conviction that every feature value is right. More resources Ready to level up your underwriting models Get started with Chalk today and build fraud detection infrastructure that moves at the speed of your business. Explore additional use cases # Build real-time ML systems that drive growth source: https://chalk.ai/use-case/growth-decisioning Build real-time ML decisioning systems for retention, pricing, and personalization using live customer data across the entire lifecycle. TALK TO AN ENGINEER GROWTH DECISIONING Chalk helps data and ML teams power real-time decisioning systems that personalize, predict, and optimize revenue across the customer lifecycle. Trusted by leading revenue intelligence and growth teams MoneyLion powers dynamic product offers and retention models with Chalk, computing user and account features continuously and serving them to models in milliseconds. Read story Growth decisions in real-time More resources Real-Time ML for Marketing, Retention, and Pricing Get started with Chalk today and transform your ML workflows. Explore additional use cases # Ship agents that work. source: https://chalk.ai/workflow/evaluate-agents Replay agents against the state your systems actually served, score every run against the trace, and use the traces to make the model better. Prove the change is better before shipping. All in your cloud. Read the Docs TALK TO AN ENGINEER EVALUATE & IMPROVE AGENTS Simulate real-world agent behavior, evaluate changes before you ship, and monitor every run after launch. Teams running AI/ML in production with Chalk Recreate exactly what an agent would have seen at a specific point in time, including its data, tools, and conditions. Chalk controls tool access through a secure gateway, so you can enforce knowledge cutoffs and run a historical scenario at scale. Simulate point-in-time, real-world environments. Score outcomes, compare versions, and inspect traces to understand exactly what changed when you iterate on prompts, models, tools, and context. Identify whether changes make the agent better. Evaluate agents before you ship. Inspect the complete execution path, including tool calls, queries, timing, and outputs. Catch failures and unexpected behavior, and then improve your agent fast. Monitor every agent run in production. Powering the continuous loop of agent development. Chalk gives your team a complete platform for taking AI agents from idea to production. Run agent tasks in isolated, gVisor-hardened environments with their own filesystem, network namespace, and resource limits. Use sandboxes for agent work, like generated-code execution or data analysis. Sandboxes See the complete path an agent took to complete a task. Expand any run to inspect tool calls, queries, operations, timing, and outputs. Identify slow or failed steps quickly. Agent traces Define evaluations. Run large evaluation sets in parallel. Score results with Python or an LLM judge, compare models and versions, and drill from a high-level score into an individual agent trajectory. Evaluations, scorers, and runs Generate structured, time-consistent context for historical evaluation. Create versioned evaluation datasets from your historical data. Use offline queries to recreate the context an agent had at a specific point in time. Then run repeatable evaluations against the same scenarios whenever you change your agent. Context Engine Evaluate and improve agents on the AI data platform for inference. Go deeper on agent evaluation. How a single knowledge cutoff locks an entire agent trajectory to a point in time. Read the announcement → BLOG /blog/chalk-compute Introducing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud. DOCS https://docs.chalk.ai/docs/compute/sandbox Get started with sandboxes Building an agent is easy. Deploying it confidently is hard. Chalk was the only integrated compute and context engine for agents… What would have taken months took weeks. Bring us the agent proof-of-concept that needs to perform reliably in production. We'll help you set up an evaluation harness that gives you the confidence to deploy your agent to production. # Host AI models without building an inference platform from scratch source: https://chalk.ai/workflow/open-model-serving Run open-weight models like GLM-5.3, Deepseek v4 Pro, or Kimi K3 while keeping inference and data within your environment. Reduce the cost of model calls compared to running frontier closed state-of-the-art models and maintain data sovereignty. Read the Docs TALK TO AN ENGINEER SERVE OPEN & CUSTOM MODELS IN YOUR CLOUD Trusted by teams running AI/ML in production Avoid dependency on an inference infrastructure provider, and meet performance, reliability, and cost targets without undifferentiated platform work. No per token markup, and inference is next to your data. Control your inference costs. Inference, model weights, prompts, outputs, and traffic remain in your cloud environment. Run production AI in your own cloud, not somewhere else. Data doesn’t leave your cloud. Own your data. Automatically route each task to the right model based on its complexity, optimizing for the best balance of cost and speed. Manage and route inference across models. One-click deploys for open models. Chalk handles the orchestration, so you can launch models, scale with demand, and operate inference easily while retaining your control. Run inference workloads on serverless or self-hosted infrastructure. Chalk handles the deployment so you don’t have to worry about infrastructure. Compute Scale to zero when there is no traffic so you only pay for capacity when you're using it. When traffic returns, Chalk provisions capacity and gets the model serving automatically. Chalk handles optimized model weights fetches so a new replica goes from zero to serving fast, and KV Cache aware routing to route follow-up requests to the same replica with the cache for a given conversation. Scaling Groups Distributed file system optimized for model weight distribution. Store the weights for your model rather than downloading them from Hugging Face every time. Download the weights into a volume, and then load them fast from there to keep weights inside your VPC, reduce dependency on external services, and speed up model startup. Volumes Create keys, and track total token requests and usage by key. Set daily budgets, and block if a limit is exceeded. AI Router Call deployed models from online or offline Chalk queries, and use production data to build datasets for supervised fine-tuning, all within the same cloud environment where you serve your models. Context Engine Run inference, model weights, prompts, outputs, and traffic in your own cloud environment, without taking on the operational burden of deploying and scaling the infrastructure yourself. Your cloud Serve models on the AI data platform for inference. Go deeper on model serving. Learn More https://docs.chalk.ai/docs/compute/running-your-own-models Run any HuggingFace model with vLLM - and Chalk supplies the GPU, the container, and the URL. Docs https://docs.chalk.ai/docs/compute/overview Learn about Chalk Compute DOCS https://docs.chalk.ai/docs/compute/scaling-groups Get started with scaling groups Bring us your most expensive monthly model spend bill. We will fine-tune an open model on your data, help you serve it in your cloud and compare quality and cost side by side. # Turn your proprietary data into a model you own. source: https://chalk.ai/workflow/fine-tune-models Train and fine-tune open models on your proprietary data with Chalk. Build point-in-time-correct datasets, evaluate models side by side, and deploy the best model into production in your own cloud. Read the Docs TALK TO AN ENGINEER TRAIN & FINE-TUNE OPEN MODELS Fine-tune open models on your unique data to build models that deliver performance at a lower cost. Teams running AI/ML in production with Chalk Build point-in-time-correct datasets and fine-tune open-weight models on the examples that matter to your business. Train on your proprietary data. Test open models, fine-tuned models and closed models against the same task. Compare performance so you know which model actually performs best for your task. Evaluate models side by side. Take the model that wins and put it into production. Deploy your custom model on scalable GPU infrastructure in your cloud and serve it fast. Deploy in one click. The best model for your task is the one you train yourself. Frontier models are built to solve arbitrary tasks. Train smaller models to run faster and cheaper on your actual workload. Chalk turns proprietary data into training data, gives you the tools to prove which model performs best, then puts it into production. Build point-in-time-correct datasets that reflect your business context. Join and query data across your sources without moving or duplicating data into another platform. Context Engine Evaluate models against a determined task. Define evaluations in code or the UI, score with Python or LLM judges, and compare performance with scores. Evaluations Easily deploy the best model for your task with one click. Serve fine-tuned models on scalable GPU infrastructure and expose them through a production-ready endpoint. Compute Use your model where it works best. Route requests across custom models, open models, and frontier models so you can send the right workload to the right model, configure fallbacks, and shift more traffic to your custom model as it improves. AI Router Keep your data and models in your cloud. Your proprietary data never needs to leave your perimeter. Run models in your own cloud environment, under your existing infrastructure and access controls. Your Cloud Train and fine-tune models on the AI data platform for inference. Chalk brings data, evaluation, and inference into one platform so you can turn proprietary data into a production model that outperforms. Go deeper on model training. LEARN MORE https://docs.chalk.ai/docs/compute/model-inference Run Hugging Face models with vLLM on your own GPUs DOCS https://docs.chalk.ai/docs/compute/overview Learn about Chalk Compute https://docs.chalk.ai/docs/compute/scaling-groups Get started with scaling groups Your data is your model's advantage. Cut your model bill. We'll fine-tune an open model on your data, help you serve it, and show you exactly how it compares on performance and cost. # Run agents safely. source: https://chalk.ai/workflow/agent-sandboxes Run agents and untrusted code in gVisor-isolated sandboxes inside your cloud. Deny by default, grant per workload, and give agents tools without giving them credentials. Read the Docs TALK TO AN ENGINEER SANDBOX AGENTS & UNTRUSTED CODE Agents need to write and execute code, work with sensitive data, and interact with systems outside the model — safely. Chalk gives every agent an isolated sandbox that runs inside your cloud with the identity, network controls, data access, and tools you define. Teams running AI/ML in production with Chalk Run inside the environment you already govern. Chalk's agent runtime deploys sandboxes directly inside your private cloud. InfoSec clears Chalk once inside your VPC, allowing your team to ship new agent use cases without new review cycles. In your cloud. Run in gVisor-isolated sandboxes with their own filesystem, network namespace, resource limits, and CPU/GPU allocation. Least privilege is the starting state. Add outbound access by hostname, CIDR, and port ranges allowlist. Each sandbox launches with its own OIDC-compliant cloud identity, scoped to that workload alone. Secure. Start isolated sandboxes in under a second, and scale up to thousands of concurrent workloads. From a single coding agent to large-scale parallel agent jobs, Chalk handles the execution layer. Fast, at scale. Agents that your CISO approves. Designed to help AI teams deploy safely and securely. Production-grade by default. Chalk Compute was the only integrated compute and context engine for agents that ran entirely inside our own environment, at the scale we needed, without becoming an infrastructure project. What would have taken months took weeks. Run every agent in a gVisor-isolated sandbox with its own filesystem, network namespace, and resource limits, with support for CPU and GPU workloads. Control outbound access with hostname, port, and CIDR allowlists. Define which services and IP ranges an agent can reach, while connections to non-allowlisted destinations are dropped at the network layer. Sandboxes and network policy. Give each sandbox a workload-scoped cloud identity, rather than embedding long-lived credentials in agent code or images. Let agents hit endpoints with a dummy token, and the sandbox injects the right authentication token. Workload proxying and identity. Register approved tools with the MCP Gateway, authenticate each sandbox through its workload identity, and route tool calls through a governed proxy for centralized policy enforcement and auditability. Enforce Rego policies on every MCP tool call, making decisions based on the user, scopes, backend, tool, and tool-call arguments. MCP Gateway. Mount a repo or a dataset as a versioned volume and parallelize work with no copies. Deploy Python callables as remote endpoints. Run durable agent loops and bulk jobs through a function queue with retries. Volumes and functions that scale. The same feature definitions that serve models in production provide the context for agents. Context Engine. Run agents safely on the AI data platform for inference. Go deeper on running agents safely. Why agent execution belongs next to the data, and what a generic container runtime cannot give you. Read the announcement → Blog Introducing Chalk Compute: Time-Traveling Agent Sandboxes in Your Cloud. DOCS https://docs.chalk.ai/docs/compute/sandbox Get started with sandboxes https://docs.chalk.ai/docs/security Security architecture Building an agent is easy. Deploying it safely is harder. Bring us the agent deployment that's stuck in security review. We'll show you how to deploy agents with complex data and governance needs in your cloud, at massive scale.