# Homepage source: https://chalk.ai/ Chalk | The data platform for AI + ML Chalk is the real-time data platform for ML & GenAI. More than a feature store, Chalk delivers fresh features in under 5ms with on-demand compute and caching in your cloud. The data platform for AI + ML Ultra-fast data pipelines, caching, on-demand compute. All in your cloud. Announcing our $50M Series A Customers - Ramp - Melio - Whatnot - Socure - Vital - Moneylion - Apartment list - Found - Pipe - Turo Feature pipelines in idiomatic Python Powerful data engineering workflows, without the infrastructure headaches. Powered by Rust. Built-in scheduling, streaming + caching Complex streaming, scheduling, and caching, all defined in simple, composable Python. Composed & queried in real-time Make ETL a thing of the past. Fetch all of your data in real-time, no matter how complex. Toolchain for LLMs Incorporate deep learning and LLMs into decisions alongside structured business data. Machine learning infrastructure is painful. Chalk makes it simple for data teams to focus on building the unique products and models that make their businesses thrive. Customer Testimonials Our competitive edge hinges on the speed at which we analyze risk patterns, test new detection features, and deploy updates. With Chalk, our iteration speed went from days to hours. - Raine Scott, Co-Founder, Verisoul We fully replaced our legacy URL candidate suspiciousness scoring model with a more powerful in-house one, shipped in record time! We're optimistic our feature ideation-to-production cycles will only get faster as we 10x our scale this quarter. Huge thanks to the Chalk team for your amazing support—this wouldn't have happened so quickly without them! - Justin D'Souza, Machine Learning Engineering, Doppel Emergency room decisions can mean life or death for patients. Chalk powers these most critical predictions for us. We evaluated a wide range of technical solutions and Chalk stood out for having the very best developer experience at the most competitive cost. - Te Riu Warren, CTO, Vital Chalk has become a critical component of our Risk Intelligence Platform. It expanded Ramp's capabilities with online machine learning and enabled us to scale safely by powering our transaction fraud model and credit underwriting process. - Ryan Delgado, Director of Engineering, Data Platform, Ramp Deploy to your own infrastructure: Use your existing database as your online + offline store. No bespoke storage. Everything in your cloud. High-volume workloads at ultra-low latency: Chalk's Compute engine scales horizontally out-of-the-box and executes most complex queries on a Rust-based runtime for maximum performance. 100,000 QPS in < 5ms? We have you covered. Power real-time decisions with real-time data. Make better predictions with fresher data. Don't pay vendors to pre-fetch data you don't use. Query data just-in-time for online predictions. Perfect auditability. Know everything you computed and data replay anything. Unify training and serving. Iterate faster. Experiment in Jupyter, then deploy to production. Prevent train-serve skew and create new data workflows in milliseconds. Detect, troubleshoot, and eliminate data issues. Track data use, drift, and quality effortlessly with observability—built right in. Integrations Integrate with the tools you already use and deploy to your own infrastructure. - PostgreSQL - Snowflake - AWS - Google Cloud - Databricks - Jupyter - Datadog - PagerDuty - Slack - Python - Apache Arrow - Apache Airflow Start building with Chalk: https://chalk.ai/book-demo Helpful homepage links - Book a demo: https://chalk.ai/book-demo - Documentation: https://docs.chalk.ai/docs/what-is-chalk - Code Examples: https://chalk.ai/code-examples - Chalk Blog: https://chalk.ai/blog - Customer stories: https://chalk.ai/customers - Careers at Chalk: https://chalk.ai/careers - About Chalk: https://chalk.ai/about # About source: https://chalk.ai/about Data infrastructure to answer the world's hardest questions Chalk's data platform provides the essential building blocks for machine learning with an experience developers love. Why Chalk? Solving large-scale ML problems is a difficult problem that requires mathematical precision and a lot of hard work. Our name is inspired by the favorite tool of mathematicians around the world—the chalkboard. Our logo is taken from the QED symbol, which indicates the end of a mathematical proof. Founders With a combined 35 years of experience working in industry, our founders have solved some of the most difficult problems in finance, data infrastructure, and risk management at companies like Google, Stripe, Affirm, and Palantir. Marc Freed-Finnegan - Co-Founder & CEO Marc spent many years at Google where he helped to launch the first version of Google Wallet. He went on to start Index, which Stripe acquired as its in-store payment solution—now called Stripe Terminal. Elliot Marx - Co-Founder Elliot started his career at Affirm where he built the early risk and credit data infrastructure system (the inspiration for Chalk). He then co-founded Haven Money, which Credit Karma acquired to power its banking products. Andy Moreland - Co-Founder Andy worked at Palantir on large government data infrastructure projects. He then co-founded Haven Money (with Elliot), which now powers Credit Karma Money. Careers Work with some of the world's top engineers on cutting-edge problems in systems engineering, query planning and optimization, and distributed analytical data processing systems. We are a diverse group of people working to change the world of machine learning and bring powerful new tools to the industry. # Careers at Chalk source: https://chalk.ai/careers Careers at Chalk Work with some of the world's top talent on cutting-edge problems. We're creating the data platform that fundamentally changes what's possible for developers. Why our team loves working at Chalk Working here pushes you to master systems programming at a speed you can't get anywhere else, because the problems we tackle here have a theoretical bent that's genuinely hard to find in corporate software jobs. This is where you come to level up. - Sai Atmakuri, Software Engineer When I first came to the Chalk office for my onsite, I saw an interesting problem scribbled on the chalkboard. When I asked how they solved it and they mentioned bloom filters, I knew right away that these were the people I wanted to build with. - Bill Qin, Software Engineer What began as wild whiteboarding sessions in our NY office has evolved into a real growth engine. We're building a GTM team that's sharp, creative, while doubling down on what works. If you want to help shape how truly groundbreaking AI reaches the market, this is the place to be. - Alexandra Kane, VP of Revenue I think being around sharp, driven people naturally elevates your own work. It's no surprise we have chess IMs and competitive Smash players on the team - people here bring that same intensity to everything they do. It makes our weekly board game nights pretty memorable too. - Kelvin Lu, Software Engineer If I had to describe Chalk in a nutshell, it's nice nerds solving interesting problems together. One of my favorite activities here is the biweekly book club on AI/ML - it's refreshing to collaborate with people who are brilliant but also down-to-earth and fun to be around. - Melanie Chen, Forward Deployed Engineer It's cool to get a front row seat to a company powering the backbone of one of the most exciting industries in the world right now. But what truly keeps me here is the culture of ownership. No BS, no office politics, just the freedom to move fast and see your impact immediately. - Brian Sussman, Operations Manager Life at Chalk You'll find us celebrating wins together — whether that's solving complex problems or closing important deals — while battling in board games and going on culinary adventures to fuel our energy. Benefits you'll receive - Full medical, dental, and vision insurance - Flexible Spending Account (FSA) and Health Savings Account (HSA) - Expert healthcare guidance and live one-on-one support - 401(k) retirement savings - Generous holiday schedule - Generous PTO annually - Daily lunch and dinner on us - Flex commuter benefits - Grab a ride home on us # Capabilities - Feature Store source: https://chalk.ai/capabilities/feature-store Feature Store | Every feature has a single source of truth Chalk's feature store powers production ML. Define features once, and Chalk computes them on demand for training, batch scoring, and inference. Features stay fresh and consistent across environments. One source of truth Traditionally, teams write one pipeline to fetch data for training and another to fetch that same data for production. To make matters worse, teams often re-write pipelines for each model that relies on that data which can lead to significant bugs. Chalk solves this with a unified feature repository — a single pipeline to fetch data which is then accessible for all training and production models. Hot-reload your pipelines Developing new features or iterating on existing ones can be a slow and painful process. With Chalk, preview feature updates in real time, so you can iterate quickly and develop better pipelines faster. Chalk even backfills features using historical data so you can immediately start using it to train models. Time-Travel When computing historical training data, it's common to accidentally include data "from the future." Chalk's time-travel functionality makes it easy to compute historical feature sets that accurately show how your features would have appeared in the past. Low latency serving Chalk's feature serving platform scales horizontally out-of-the-box and can handle high-volume production workloads at low-latency. Feature Discovery Stop reinventing the wheel for each new model. Chalk's feature catalog makes it easy to discover and organize the features that your team develops. Version and track feature definitions and metadata in code stored in git for easy review and change control. Audit Chalk automatically tracks detailed metadata about the provenance of each feature computed and served to your models. This makes it easy to debug issues, justify decisions to regulators, and perform exploratory data analysis. # Capabilities - Compute source: https://chalk.ai/capabilities/compute Compute | Time traveling agent sandboxes Evaluate agents with knowledge cut-offs. Generate temporally-consistent context in your cloud. Why Chalk Compute Designed to help AI teams deploy safely and securely. Production-grade by default. Fast. Sub-second cold starts on every workload. Chalk maintains a content-addressed image cache on every node in your cluster. Once an image lands on a host, every subsequent sandbox using it boots in under a second. Common base images are pre-warmed across the fleet, so cold starts are an exception, not a default. Safe. gVisor isolation, OIDC identity, policy-bound egress. Every workload runs inside a gVisor-hardened sandbox — a user-space kernel intercepts syscalls before they reach the host. Each sandbox launches with its own OIDC-compliant cloud identity, scoped to that workload alone. Outbound traffic is restricted to a hostname or CIDR allowlist you control. Yours. Runs inside your AWS, GCP, or Azure account. Always. Sandboxes execute in your VPC, on nodes you provisioned, under IAM roles you control. Customer data never crosses to a third-party plane. Logs stay in your account. KMS keys, network policies, and audit trails are the ones your security team already operates. Built for production. Agents you trust. Use cases: Financial analysis, Customer support, Document analysis, Underwriting, Loan pre-screening, Fraud triage, Sourcing, Deep research, Sentiment analysis, Coding, Data analysis. One platform. Every Agent. Compute is one piece of the unified platform, alongside Feature Store, Model Platform, Real-Time Serving, and LLM Toolchain. It all runs in your cloud. The Primitives Simple primitives. Endless possibilities. Core Infrastructure Containers — A high-level managed container that bundles image build, file upload, and lifecycle in a single Python class. Configure CPU, memory, GPU, secrets, volumes, and lifetime. Sandboxes — gVisor-isolated execution environments with their own filesystem, network namespace, and resource limits. CPU and GPU workloads, kernel-level multi-tenancy. Scaling Groups — Long-running, autoscaling HTTP services deployed from an image. Configure min and max replicas, target CPU utilization, and a port — Chalk handles the load-balanced endpoint and graceful scale-down. Filesystem Images — A fluent Python builder for defining software environments. Content-addressed caching means rebuilds are near-instant and identical across every sandbox. Volumes — Persistent, versioned file storage with copy-on-write semantics. Fork a volume to fan out parallel workloads against the same snapshot — no copies, no drift. Functions Functions — Deploy any Python callable as a remotely invocable endpoint. Automatic scaling, no servers to provision, no Dockerfiles to maintain. Function Queue — Durable, ordered task execution with retries and backpressure. The right substrate for agent loops, evals, and bulk inference jobs that have to actually finish. Bring your own agent harness. Compatible with OpenAI Agent SDK, Anthropic SDK, LangChain, LlamaIndex, and MCP Servers. Chalk Compute is the infrastructure underneath, regardless of which harness you build on. Security Hardened at the kernel. Locked at the network. Identified at the workload. gVisor isolation — Every workload runs under gVisor. Inside the sandbox, CapEff and CapBnd are zero, no-new-privs is on, and securebits are locked. Root inside the container has nothing to escalate to. Kernel surface unreachable — io_uring, bpf, perf_event_open, userfaultfd, fanotify, and kexec_load all return permission denied. /dev/kcore, /dev/mem, and /dev/port don't exist. Workload Identity Federation — Every sandbox launches with its own OIDC-compliant cloud identity, scoped to that workload alone. Network Policy — Restrict outbound egress to a hostname or CIDR allowlist; off-list traffic is dropped silently. Raw packet sockets and link-level admin are blocked outright. MCP Gateway — Sandboxes call MCP servers through the gateway, authenticated by their workload identity. Agents get tool access without ever seeing the upstream key. WireGuard Tunnels — Per-session WireGuard tunnels with dynamically negotiated keys, scoped to a single workload. Connect sandboxes to each other or bridge to on-prem databases without public internet exposure. # Capabilities - Real-Time Serving source: https://chalk.ai/capabilities/real-time-serving Real-Time Serving | Real-time feature serving for production ML Serve production features at inference time with single-digit millisecond latency. Chalk's query execution engine runs Python functions, queries databases, and calls APIs in real-time, enabling decisions on the freshest data without brittle ETL or stale caches. API & Data Integration Easily connect APIs (3rd party clients) and incorporate unstructured data with LLMs. Chalk handles auth, retries, and caching automatically. Just-in-Time Fetching Get fresh data only when needed. Chalk fetches inputs at runtime for accurate, cost-efficient predictions. Declarative Pipelines Define dependencies with Python signatures. Chalk auto-orchestrates resolvers into efficient query plans across online and offline environments. Preview Deployments Test changes in isolated preview environments. Chalk spins up sandboxes per branch for safe iteration and review. Rust-Powered Runtime Run Python at native speed. Chalk uses Rust to parallelize fetches, push down ops, and multithread computations. Built-in Observability Trace every query, monitor latency, and debug at the feature level. Chalk captures lineage and telemetry by default—no extra setup required. One language, one system. With Chalk, the same code powers training, evaluation, and inference—ensuring consistency, correctness, and eliminating the need to rewrite features. Dynamically builds efficient query plans, never fetching anything extra. Parses and transpiles your logic into static expressions that run natively. Centralizes your ML models into code, establishing a single source of truth. # Capabilities - Temporal Aggregations source: https://chalk.ai/capabilities/temporal-aggregations Temporal Aggregations | Temporal context at scale Define aggregations once and reuse them across batch, online, and real time. Chalk makes it easy to define rolling windows, decays, and normalized features directly in Python. Temporal aggregations are automatically computed and kept consistent whether you're generating training datasets, running batch scoring, or serving features at inference time. Simplify temporal feature engineering Define rolling windows, point-in-time aggregations, and decay functions once in Python. Chalk handles computation across online and offline environments without duplication. Consistent across environments The same aggregation logic that powers your training datasets also serves your production models — eliminating skew and ensuring your models behave as expected in the real world. Reusable across models Define an aggregation once and share it across every model that needs it. No copy-pasting pipeline logic or maintaining divergent implementations per team. # Capabilities - Training Data source: https://chalk.ai/capabilities/training-data Training Data | Train models with production features Generate point-in-time correct training datasets from the same feature definitions you use in production and serve production features from your training data. Keep training and inference aligned and eliminate skew. Re-use your online pipelines Instead of building an alternate pipeline for training set generation, Chalk allows you to automatically re-use online serving infrastructure to compute historical datasets. Chalk transforms inference pipelines into efficient batch pipelines and automatically time-filters data to make historical accuracy easy. Batch and Streaming In addition to online data sources, Chalk supports batch and streaming data sources. Chalk can automatically swap to data warehouses instead of online APIs or databases in order to source historical data points. Notebook Support Chalk's Python SDK works in the Jupyter notebook of your choice — local Jupyter, Google Colab, Deepnote, Hex, or Databricks — if it can execute Python, you can generate dataframes of training data. Dataset Governance Every dataset that you generate is automatically versioned and retained, which lets you seamlessly travel back in time to view the output of any past computation. Datasets can be named and shared so that teammates can use your work. Track which features are used in which datasets to help with discovery. # Capabilities - LLM Toolchain source: https://chalk.ai/capabilities/llm-toolchain LLM Toolchain | Real-time context for LLM inference A unified interface to serve, evaluate, and optimize LLMs using structured and unstructured data with sub-5ms latency. Prompt Engineering Experiment with prompts on historical data using branches. Chalk tracks outputs, computes metrics, and promotes winning prompts with one command. Model Inference Deploy inference pipelines with autoscaling and GPU support. Write pre/post-processing in Python—Chalk handles the rest, including data logging and versioning. Evaluations Log and compare model outputs with quality metrics to pick the best prompt, embedding, or model—all versioned automatically in Chalk. Embedding Functions Use any embedding model with one line of code. Chalk handles batching, caching, and lets you safely test new models on all your data. Vector Search Run nearest-neighbor search directly in your feature pipeline. Use any feature as the query, and generate new features from search results. Large File Support Process and embed large files—docs, images, videos—at scale. Chalk handles batching, autoscaling, and execution with a fast Rust backend. Real-time feature serving for LLMs Connect your LLMs to the freshest data without ETL pipelines. Retrieve structured features dynamically at inference time. Use Python (not DSLs) to define feature logic. Fetch real-time context windows with point-in-time correctness. Mix embeddings and features for fully grounded RAG workflows. Prompt engineering & evaluation Design prompts like you design software. Write, version, and reuse prompts with structured parameters. Evaluate prompts and models using historical production data. Compare model performance on accuracy, latency, and token usage. Debug failures with end-to-end traceability and lineage. Deploy prompt + model bundles as artifacts with full observability. # Use Case - Identity Fraud source: https://chalk.ai/use-case/identity-fraud Identity Fraud | Trust every identity decision with real-time fraud signals Chalk computes and serves identity, device, behavioral, and KYC features in real time. Risk models can verify users, prevent account abuse, and stop synthetic identities in single-digit milliseconds. Your Business is Specialized Off-the-shelf models only see part of the picture, so good users get blocked, and bad users get through. The best fraud teams leverage their own business insights to spot and block fraud. Chalk enables your fraud fighters and data scientists to incorporate 3rd-party signals alongside product usage, messaging and support history, and even password reset data to make high quality decisions for your unique user base. Fraud Signals are Expensive Chalk makes it easy to fetch fraud data only when it's needed. Each model specifies exactly the data staleness that it can tolerate to give you fresh data cheaply. By employing a layered approach to fraud, you can reject bad candidates without fetching expensive data. Backtest Fraud with Previews With Chalk's Branch Deployments, it's easy to experiment with new signals before going to production. Branch Deployments show how proposed features would have impacted your previous decisions. Easily Incorporate Fraud Vendors With Chalk, ship production-quality integrations on proof-of-concept timelines. Providers like SentiLink, Socure, Emailage, Whitepages, and Early Warning Systems help you form a complete picture of each user and transaction. Define Once, Use Everywhere Chalk solves fragmented pipeline logic with a unified feature catalog — a single pipeline to fetch data which is then accessible for all training and production models. Detect Fraud Early Chalk integrates with alerting systems like Pagerduty and Slack to keep your team informed about issues. Configure alerting thresholds for when data distributions don't match your expectations. # Use Case - Payments source: https://chalk.ai/use-case/payments Payments | Real-time feature computation for payments Chalk computes feature values at authorization time so models can evaluate risk using fresh, decision-time data. Just-In-Time Querying Chalk makes it easy to fetch payment and risk data only when it's needed. Each model specifies exactly the data staleness it can tolerate to give you fresh data cheaply. By employing a layered approach, you can reject bad candidates without fetching expensive data. Low-Latency Authorization Chalk's compute engine scales horizontally out-of-the-box and executes queries on a Rust-based runtime for maximum performance — scoring transactions using real-time signals across users, devices, and accounts at single-digit millisecond latency. Define Once, Use Everywhere A single pipeline to fetch data, accessible for all training and production models. No duplicated logic across teams, no discrepancies between pipeline code. Audit & Versioning Chalk perfectly tracks decisions from raw data sources to the feature values that power models, empowering teams to deeply understand the provenance of each authorization decision. Monitor Easily differentiate between organic shifts and unexpected format changes in upstream data sources, or development mistakes. Get alerted automatically before issues impact authorization accuracy. # Use Case - Underwriting source: https://chalk.ai/use-case/underwriting Underwriting | Underwriting demands real-time data Underwriting decisions depend on fresh, high-conviction data. Chalk is the federated query execution engine behind modern credit and underwriting teams. It computes features at decision time, enables faster iteration, and scores applications in single-digit milliseconds at scale. Unique Credit Decisions Your models should be as personal as your customers. With Chalk, your credit analysis is customized to your requirements, and can change as fast as your business. Empower your teams with the freedom to solve the particular problems that are most important to your business. Backtest Credit Data With Preview Deployments, Chalk allows developers and data scientists to test out their theories alongside your deployed application without polluting your on-going underwriting. Time-Travel Use Chalk's time-travel functionality to backfill new and updated features so that you can see how they would have impacted past decisions. When you like what you see, launch the changes to production immediately. Just-In-Time Querying Chalk makes it easy to fetch credit data only when it's needed. Each model specifies exactly the data staleness that it can tolerate to give you fresh data cheaply. By employing a layered approach to credit, you can reject bad candidates without fetching expensive data. Diverse Data Sources The best teams incorporate alternative data like Plaid transactions, Rutter accounting data, merchant data, and product usage to build a comprehensive risk profile. With Chalk, engineering teams can add robust integrations as fast as your data science teams propose them. Audit & Versioning Chalk perfectly tracks decisions from raw data sources to the feature values that power models, empowering teams to deeply understand the provenance of each decision — critical for regulatory compliance. Monitor Easily differentiate between organic shifts and unexpected format changes in upstream data sources, or development mistakes. Get alerted automatically before issues cause defaults. # Use Case - Recommender Systems source: https://chalk.ai/use-case/recommender-systems Recommender Systems | Power recommender systems with real-time features Chalk powers personalized feeds by computing fresh user, item, and marketplace features at request time. Teams use Chalk to serve recommendations that reflect real-time behavior and semantic similarity across dynamic inventories. Serve Personalized Feeds in <5ms Personalization depends on data freshness. Chalk serves real-time user, inventory, and session features at inference time with sub-5ms latency — powering feeds and recommendations that respond instantly to platform changes. Unify Behavioral and Inventory Data for ML Chalk unifies session data, product metadata, behavioral signals, and third-party sources into a single feature set. Whether your data comes from streaming platforms, databases, or APIs, Chalk ensures everything is queryable, real-time, and model-ready. Embeddings and Vector Search Run nearest-neighbor search directly in your feature pipeline. Use any feature as a query to generate new features from search results, powering semantic similarity and content-based recommendations. Test New Recommendation Strategies Without Breaking Production With Chalk's branching and historical backfills, data teams can test new ranking and personalization logic in isolation — with built-in version control, branch deployments, and full observability. # Use Case - Search and Ranking source: https://chalk.ai/use-case/search-and-ranking Search and Ranking | Power search and ranking with real-time features Chalk powers search models by computing relevance features and embeddings at query time. Teams leverage Chalk to adapt results instantly based on user context, behavior, and intent, while maintaining full control over ranking logic and data. Compute Query, User, and Inventory Features at Inference Time Chalk's query execution engine runs Python functions, queries databases, and calls APIs in real time — delivering low-latency feature serving so ranking models always operate on the freshest signals. Embeddings and Semantic Similarity Use any embedding model with one line of code. Chalk handles batching, caching, and lets you run nearest-neighbor search directly in your feature pipeline to power semantic search and ranking. Test New Ranking Strategies Without Breaking Production With Chalk's branching and historical backfills, data teams can test new ranking logic, promotions, and personalization in isolation — with built-in version control, branch deployments, and full observability. Point-in-Time Correct Training Data Generate training datasets directly from production feature definitions so ranking models learn from the same data they'll see in production — eliminating skew and improving model reliability. # Use Case - Growth Decisioning source: https://chalk.ai/use-case/growth-decisioning Growth Decisioning | Build real-time ML systems that drive growth Chalk helps data and ML teams power real-time decisioning systems that personalize, predict, and optimize revenue across the customer lifecycle. Real-Time Personalization Chalk serves real-time user, behavioral, and contextual features at inference time with sub-5ms latency — powering personalized experiences that respond instantly to customer actions. Unify Data Across the Customer Lifecycle Chalk unifies session data, transaction history, behavioral signals, and third-party sources into a single feature set. Whether your data comes from streaming platforms, databases, or APIs, Chalk ensures everything is queryable, real-time, and model-ready. Faster Experimentation With Chalk's branching and historical backfills, data teams can test new decisioning logic in isolation — with built-in version control, branch deployments, and full observability — before promoting changes to production. Define Once, Use Everywhere A single feature pipeline powers training, batch scoring, and real-time inference. No duplicated logic, no drift between environments, no inconsistencies across teams. # For - Data Engineers source: https://chalk.ai/for/data-engineers Chalk for Data Engineers Chalk unifies your data schema to compute and serve features defined in Python consistently across batch, training, and real time. Why data engineers choose Chalk Eliminate manual pipelines Chalk automatically builds and executes DAGs for computing your queries, removing custom orchestration. Unified feature catalog Ensure consistent feature definitions across backfills, training sets, and real-time inference from a single source of truth. Simplify temporal aggregations Define rolling windows, decays, and normalized features directly in Python. Serve in real time without extra systems Low-latency APIs without Kafka, Flink, or custom streaming stacks. Generate reproducible training data Create point-in-time-correct datasets with full lineage and versioning. Helpful links - DAG orchestration: https://docs.chalk.ai/docs/architecture#data-orchestration - Feature definitions: https://docs.chalk.ai/docs/resolver-online-offline - Rolling windows: https://docs.chalk.ai/docs/materialized_aggregations - Low-latency integrations: https://docs.chalk.ai/docs/integrations - Point-in-time correctness: https://docs.chalk.ai/docs/temporal-consistency # For - MLOps source: https://chalk.ai/for/mlops Chalk for MLOps Build fully observable ML pipelines with automatic data lineage on every query. Why MLOps engineers choose Chalk Data lineage and traceability Full visibility into how every feature is built for better debugging. Built-in versioning and auditability Automatic versioning and audit trails for consistent, reproducible feature definitions. Feature monitoring and alerting Monitoring with alerts so features remain fresh, accurate, and drift-free. Infrastructure observability Logs for debugging, Kubernetes cluster activity, and system-level metrics with alerts. On-demand pre-computation On-demand pre-computation that replaces brittle pipeline logic and manual orchestration. Low-latency serving at scale Sub-5ms real-time serving at high throughput. Helpful links - Data lineage: https://docs.chalk.ai/docs/feature-discovery#data-lineage - Versioning: https://docs.chalk.ai/docs/feature-versions - Monitoring: https://docs.chalk.ai/docs/metricmonitor - Alerts: https://docs.chalk.ai/docs/alertconfig - Kubernetes observability: https://docs.chalk.ai/docs/kube-resources-drilldown - Pre-computation: https://docs.chalk.ai/docs/feature-caching - Architecture: https://docs.chalk.ai/docs/architecture # For - ML Engineers source: https://chalk.ai/for/ml-engineers Chalk for AI + ML Engineers Experiment faster, unify structured and unstructured data, serve features in real time, and power models with production-grade infrastructure, all from a single Python interface. Why AI/ML engineers choose Chalk Unified data for every modality Work with structured data like transactions and aggregations, and unstructured inputs like text, embeddings, and prompts. Faster experimentation loop Experiment faster by defining features once in Python and reusing them across training and inference. Reliable, point-in-time datasets Generate reliable training data with point-in-time correct backfills. Low-latency model serving Serve models in real time with sub-5ms feature retrieval from heterogeneous sources. LLM-native feature workflows Power LLMs with embeddings, vector search, and prompt construction. Reproducible by default Reproduce and monitor all queries with built-in versioning and query logs. Helpful links - Branches: https://docs.chalk.ai/docs/branches - Aggregations: https://docs.chalk.ai/docs/aggregations - Embeddings: https://docs.chalk.ai/docs/embeddings - Prompts: https://docs.chalk.ai/docs/prompts - Backfills: https://docs.chalk.ai/docs/backfilling-data#backfilling-from-datasets - Sub-5ms serving: https://docs.chalk.ai/docs/architecture - Vector search: https://docs.chalk.ai/docs/vector-search - Versioning: https://docs.chalk.ai/docs/feature-versions#default-versions # For - Data Scientists source: https://chalk.ai/for/data-scientists Chalk for Data Scientists Discover and create features, generate reproducible datasets, and accelerate AI and ML experiments with Chalk. Why data scientists choose Chalk Data scientists move quickly from hypothesis to production. With Chalk they: Observability into served features Get full visibility into the features being served in production for better debugging and trust. Training datasets with lineage Build training datasets with full lineage so every experiment is traceable and reproducible. Training and production always in sync Trust that training and production feature definitions always stay aligned — no skew, no surprises. # Customers - Whatnot source: https://chalk.ai/customers/whatnot How Chalk fuels Whatnot's live shopping platform Whatnot uses Chalk to power real-time recommendations across its marketplace, from feed ranking to show discovery. By replacing a batch-based system with a unified feature platform, Whatnot delivers fresher, more personalized experiences, iterates on models faster, and scales its machine learning systems to meet the demands of its fast-growing live commerce platform. Use Case: Recommendations Industry: Marketplace Cloud: AWS Challenges: - Coverage fell to ~90% as scale and inventory increased - Personalization was delayed up to 24 hours due to cold starts - Session-level signals like taps and chats were being omitted Solutions: - Deliver low-latency computation at scale with 300M+ features/sec - Provide a unified feature store and real-time engine with Chalk - Enable faster deployment of models with fresher, live marketplace features Whatnot is the largest live shopping platform in the U.S. and Europe, a marketplace where buyers and sellers connect through real-time, community-driven commerce. The company surpassed $3 billion in live sales last year alone, with users spending on average 80 minutes a day browsing and buying from live shows. The Challenge: As Whatnot scaled, its ML systems started to lag behind product needs. The original batch-based architecture, which generated billions of user–show predictions nightly, struggled to keep up with fast-changing inventory, new seller launches, and real-time user behavior. Cold starts were common, coverage dropped, and model iteration slowed — all of which impacted recommendation quality. The Solution: Whatnot adopted Chalk to power its real-time feature infrastructure, serving as both the feature store and online feature engine. Feature logic is defined once in code and reused across batch and online systems. At inference time, Chalk materializes features from request context and historical data. Payloads exceed 1MB and include tens of thousands of features, yet the system consistently delivers responses under 150 milliseconds, even during peak traffic. Outcomes: - Cold start delay reduced from ~24 hours to under 1 hour - Prediction latency reduced from ~24 hours to ~150ms end-to-end - Dev iteration time reduced from weeks to days - Personalization reach improved from ~90% to 99.9% of users receiving a real-time personalized feed - System supports 300M+ features/sec with >99.99% uptime - Load tested to handle 9x future traffic # Customers - Turo source: https://chalk.ai/customers/turo How Turo Built a Self-Serve ML Feature Platform for Search and Pricing with Chalk Turo standardized feature delivery across search, pricing, and risk using Chalk's compute-first feature store, enabling faster iteration, predictable production performance, and a scalable foundation for real-time ML. Use Case: Search and ranking, Dynamic pricing Industry: Marketplace Cloud: AWS Challenges: - Fragmented feature pipelines across production systems - Slow time to production for new features - Inconsistent feature workflows across models Solutions: - A single, centralized feature store replacing fragmented pipelines - Self-serve feature delivery accelerating time to production - Standardized, resolver-driven feature workflows shared across models Turo is the world's largest car-sharing marketplace, with more than 300,000 host-owned vehicles across the US, Canada, the UK, France, and Australia. Machine learning systems are foundational infrastructure — determining how guests discover cars, how owners price trips, and how trust is enforced across the marketplace. The Challenge: Before Chalk, Turo did not have a unified feature store for production feature delivery. Feature data lived in multiple places. ML engineers had to comb through production databases to find viable data sources, then craft Airflow jobs and custom pipelines to make feature data available for online use cases. Feature definitions and model logic were scattered across repositories and services, making it difficult to treat features as shared, reusable production assets. The Solution: Turo adopted Chalk to create a single, repeatable path for production feature delivery, using Chalk as its centralized feature store owned by the ML engineering team and deployed in Turo's own cloud. Feature definitions live in one place and follow consistent patterns across different models. Chalk absorbed much of the feature-related data engineering work that previously required custom pipelines or cross-team coordination, covering roughly 80% of feature delivery needs through standard workflows. Outcomes: - Time to ship a new production feature reduced from 2–3 weeks to 1 week - Feature delivery moved from team-dependent coordination to self-serve within the ML engineering team - Shared resolver patterns now cover most use cases across search, pricing, and risk models - Predictable production latency that teams can plan around # Customers - Apartment List source: https://chalk.ai/customers/apartment-list How Apartment List uses Chalk to power personalized apartment recommendations Apartment List leverages Chalk's real-time feature platform to deliver an individualized and intuitive apartment search experience. By unifying fresh data from SQL, streams, Python UDFs, Expressions, and APIs in real time, Apartment List runs sub-5ms queries for its dynamic recommendations as renters refine their search. Use Case: Recommendations Industry: Marketplace Cloud: GCP Challenges: - Delays in feature computation impacted search relevance for renters updating preferences - Manually defining and fetching features led to duplicated logic, inconsistencies, and slow deployment cycles - Querying multiple sources introduced data latency and performance bottlenecks Solutions: - Enables immediate search updates for a more responsive UI - Delivers dynamic personalization for more relevant recommendations - Supports seamless integration of heterogeneous data sources through direct API calls The Challenge: Finding an apartment is an interactive and personalized process. Renters continuously refine their budget, location preferences, and desired amenities—expecting search results to adjust instantly. If preferences fail to update in real time, users encounter stale or irrelevant listings, leading to frustration and drop-off. The Solution: Apartment List implemented Chalk's feature platform, transforming how it delivers personalized, low-latency search recommendations. With Chalk, Apartment List: - Enables immediate search updates for a more responsive UI - Delivers dynamic personalization - Reduces latency for ranking Outcomes: - Real-time search personalization: Search results adjust instantly when users change preferences - Low-latency feature retrieval: Sub-5ms response times keep ranking updates highly performant - Seamless API integration: Chalk integrates with Apartment List's APIs, supporting both batch and streaming data services # Customers - Mission Lane source: https://chalk.ai/customers/mission-lane Mission Lane's Source of Truth for Credit Decisioning with Chalk Mission Lane uses Chalk to power real-time credit approvals, fraud detection and customer-facing features, all from a single, consistent feature platform. As modeling and decisioning became core to the business, Chalk gave the Mission Lane team the infrastructure to move faster, reduce drift, and scale confidently across systems. Use Case: Credit underwriting, fraud detection, in-product data Industry: Fintech Cloud: GCP Challenges: - Fragmented feature logic across Python, SQL, and production systems - Inconsistent model behavior across batch, real-time, and training pipelines - No centralized system to track, version, or reuse features across teams Solutions: - Unified feature definitions across batch, real-time, and training pipelines - Native support for Python, SQL, and hybrid cloud deployment - Centralized platform for decisioning used by ML, ops, and product teams Mission Lane is a fintech helping millions of Americans left behind by traditional financial service companies access fair, transparent credit. As modeling and decisioning systems became more complex, Mission Lane needed a modern feature platform to unify real-time and batch infrastructure. The Challenge: Across the company, different teams operated with different tools, leading to duplication, silent inconsistencies, and potential drift between training and production. The hybrid architecture required implementing the same feature logic in multiple places. The Solution: Mission Lane chose Chalk as its next-generation feature platform to unify batch and online decisioning and eliminate drift across workflows. Chalk now powers both mission-critical model evaluations and emerging use cases beyond machine learning. Outcomes: Chalk helps Mission Lane move faster, improve reliability, and extend value beyond its ML team. The platform enables faster deployment of new features, consistent training/serving, and improved model iteration velocity. # Customers - MoneyLion source: https://chalk.ai/customers/moneylion MoneyLion delivers AI-powered personal finance products with Chalk MoneyLion uses Chalk's real-time data platform to unify ML development, accelerate feature delivery, and reduce time-to-production. By centralizing experimentation, serving, and governance, MoneyLion scales AI applications across fraud prevention, budgeting, and personalization. Use Case: Fraud, Recommendations Industry: Personal Finance Cloud: AWS Challenges: - Fragmented collaboration across ML lifecycle - High engineering overhead and offline-to-online drift - Lack of centralized governance and feature reuse Solutions: - Unified ML development on shared platform - Python-native feature development for seamless experimentation-to-production - Feature store with built-in governance and versioning MoneyLion is on a mission to empower Americans to make better financial decisions. With millions of users across lending, investing, and personal finance tools, the company relies on a complex machine learning ecosystem to drive real-time fraud detection, customer engagement, and personalized recommendations. The Challenge: Building and deploying machine learning solutions across teams became challenging as the business scaled. Different workflows, goals, and metrics made collaboration slow and costly. MoneyLion needed an alignment layer for the entire ML lifecycle. The Solution: Chalk unified MoneyLion's fragmented ML workflows by providing a developer-first platform that lets each team contribute effectively without forcing a rigid process. Chalk replaced manual micro-service maintenance with dynamic feature pipelines. Outcomes: Today, with Chalk powering real-time ML infrastructure, MoneyLion can protect users with real-time fraud detection, deliver smarter budgeting and spend insights, and surface contextual feed and lifecycle recommendations. # Customers - Verisoul source: https://chalk.ai/customers/verisoul Verisoul stops fake accounts with Chalk's real-time feature platform Verisoul leverages Chalk's real-time feature platform to combat evolving fake account threats, shipping detection updates 10x faster and using fresh inference-time data for 4x more accurate fraud detection. Use Case: Detection models, risk tree decisioning Industry: Fraud detection Cloud: GCP Challenges: - Data staleness for risk decisions impacted decisioning speed and quality - Slow and inefficient signal iteration causing longer deployment - Lack of auditability created engineering and customer service pain Solutions: - Now deploys detection updates 10x faster - Real-time inference data improves detection accuracy 4x over traditional methods - Full auditability enables transparent and explainable fraud decisions for customers Verisoul specializes in stopping the most advanced fake account bots before they cause harm. The platform analyzes device signals, behavioral patterns, biometric anomalies, and network-level risks in real time. The Challenge: Before implementing Chalk, Verisoul relied on manual, in-house feature engineering pipelines. This led to data staleness, slow signal iteration, and lack of auditability. The Solution: By leveraging Chalk, Verisoul transformed its decision models into a true real-time system. Instead of relying on outdated or precomputed risk signals, they now make decisions using fresh, continuously updated data. Outcomes: Since implementing Chalk, Verisoul has significantly improved detection speed, accuracy, and scalability: - Real-time fake account prevention at scale - 10x faster feature development and deployment - Full auditability & reduced engineering overhead - Increased engineering efficiency and speed # Customers - Vital source: https://chalk.ai/customers/vital Vital predicts hospital wait times with Chalk's feature platform Using advanced AI, Vital transforms complex health record data into easy-to-use, personalized interfaces that inform and engage over one million patients per year. After deploying Chalk, Vital was able to stop wrangling infrastructure and focus on improving its models, launching new products, and delivering a world-class patient experience. Use Case: ERAdvisor: Real-time hospital wait times Industry: Healthcare Cloud: AWS Challenges: - Experimentation slow, tedious, unreliable - Feature changes required deep, specialized knowledge limited to a handful of engineers - Third-party data infrastructure Solutions: - Now update model 2-4 times / month - Up to nine engineers working concurrently - Patient data never leaves the cloud Vital is redefining patient experience with software that gives more control, clarity, and predictability to emergency department visits and hospital stays. Hospitals across the U.S. use Vital to improve patient satisfaction, drive growth and patient loyalty, achieve better clinical outcomes, and reduce workload for care teams. The Challenge: Vital's previous solution created challenges due to poor developer experience and architectural limitations. Experimentation was slow and required deep knowledge of complex systems, and data processing occurred on third-party infrastructure. The Solution: Vital selected Chalk to address these challenges, dramatically increasing the pace of product development with rapid iteration and experimentation, concurrent engineering work, and full control of customer data. Outcomes: Vital now uses Chalk to process and serve features powering seven different models in production, including ERAdvisor software which guides patients through emergency room visits with personalized wait times and next steps. The team now deploys updated models 2-4 times a month, with improved model accuracy, performance & cost efficiency, data privacy and security, and developer happiness.