Customer Story

How Grid accelerates model training and creates continuous improvement with Chalk

customer story mobile landing image

Client

Use Case

Cash advance underwriting

Industry

Fintech

Cloud

GCP

Challenges

  • Fragmented model development across data, manually managed GPUs, and MLflow
  • Training experiments required engineering handoffs to provision compute
  • Separate feature implementations for training and inference risked inconsistencies

Solutions

  • Unified data prep, training, evaluation, and serving in Chalk, cutting development time by 50%+
  • On-demand CPU/GPU training eliminated infra handoffs and enabled 3x as many runs
  • Shared feature definitions across training and inference
  • Continuous model development, with scheduled retraining

Overview

Grid is a consumer fintech company that gives members a cash advance through its Cash Advance product. Its first-advance underwriting model determines which new members to approve, using bank transaction data to predict repayment risk. With Chalk Model Training, Grid cut its model development loop by more than half, giving the team more speed to improve its underwriting models and helping it approve more qualified members at the same loss rate.

  • 50%+ faster model development: Cut the end-to-end process from data preparation through evaluation and serving by more than half.
  • 3x training velocity: Increased training runs from ~10 to ~30 per week by running models and hyperparameter settings in parallel.
  • 0 engineering handoffs: Data scientists can launch CPU and GPU training jobs without managing infrastructure.
  • Shortened time from idea to production from weeks to days: Shortened the path from a new model idea to a dry-run on live traffic.
  • One continuous workflow for experimentation: Train, compare, evaluate, and prepare models for serving in Chalk.

Grid’s first-advance model is built for the hardest approval decisions

Grid is a consumer fintech app that brings everyday financial services into one place. Its Cash Advance product gives members up to $200 between paychecks, helping them cover short-term cash gaps without an overdraft.

Every advance is an approve-or-decline decision made in real time. If Grid approves too many risky members, losses increase. If Grid declines too many good members, then Grid misses valuable customers. Better risk ranking lets Grid approve more good members without increasing losses.

The hardest decision is the first advance: a new member has no repayment history with Grid, so the model has to predict risk from the financial signals available. Most underwriting models make decisions based on bank history summary metrics, like average balance, total deposits or number of overdrafts. Grid's first-advance model looks beyond amounts and balances. It also reads transaction descriptions such as "ADP PAYROLL," "Venmo," and "DoorDash." A sentence transformer converts that transaction text into embeddings, which the model combines with amount, timing, and rolling cash-flow features to predict default risk.

“The first advance is our hardest underwriting decision because we’re making it without any repayment history. We need to get as much signal as we can from the financial data available to us. Chalk has made it much faster to test new approaches and get the strongest models into production.”
hi
Bijan Safai Software Engineering Manager

Before Chalk Model Training: a fragmented model development process

Grid’s engineers spent significant time provisioning GPUs, moving data, rebuilding features, recomputing embeddings, and connecting different tools to run and evaluate experiments.

The training setup happened outside Chalk, across a collection of tools and scripts: a training framework, cloud GPUs, a separate experiment tracker, and custom code to connect everything together. The result was a fragmented workflow where every new experiment meant rebuilding pieces of the same process.

  • Every model experiment had to wait for another team and infrastructure configuration. A Grid data scientist couldn't start an experiment without an engineer's time to simply start a GPU or training run. Training ran on a single, hand-maintained GCP GPU instance which meant the team managed drivers, CUDA and quotas themselves, pulled data from BigQuery onto the machine for every run, and paid for the GPU even when it sat idle.
  • Experiment tracking lived separately, and shipping was manual. Metrics lived in MLflow, features lived in Chalk, and comparing candidates meant stitching the two together by hand. This meant slower decisions about which model should be promoted to production. Before moving a model live, the Grid team had to find the right checkpoint, register it and export a serving version by hand.
  • Two disconnected versions of the same features. Training built features from BigQuery in its own script, while live scoring rebuilt them in separate inference code. Any mismatch between the two quietly degraded live scores, and the first sign would be an increase in approvals for unqualified members or repayment defaults weeks later.
“Training was the part of our workflow that took the most engineering effort. I was spending time waiting for another team to provision GPUs, moving data around, and stitching experiment results together instead of working on the model itself. Every experiment had infrastructure work in front of it. Now the question is just whether the idea is good.”
hi
Hyojin Lim Machine Learning Engineer

Announcing Chalk Model Training

Chalk Model Training allows you to train models on the same features you already use in production. Chalk manages all the infrastructure underneath. You can simply define the model you want to train, pull what data to use from your Chalk features, and specify what compute hardware you need. When you start a training run, Chalk starts a Chalk Sandbox in your cloud to run your model training. It connects your dataset or query, adds your secrets, streams logs back to you and saves checkpoints.

With Chalk Model Training, requesting compute resources becomes one line of config and stops being an infrastructure project. Read more in the docs here.

Chalk Model Training accelerates the path from model idea to production

Grid had one goal: spend time on better approval decisions, rather than infrastructure. With Chalk Model Training in place alongside the Chalk features and serving Grid already used, every stage of model development now runs on one platform. That covers transforming raw transaction data, serving features, training the model that makes live approve-or-decline calls, training and deploying on managed infrastructure with GPU scheduling, and running experiment analysis to close the loop.

The team now prepares data once, launches on-demand CPU or GPU training jobs, compares models in a Chalk Notebook, and moves a winning model into dry-run and production. Here’s how a new model moves from idea to production at Grid today.

Prepare the data once.

Grid defines its production features in Chalk. For transaction descriptions, Grid generates embeddings once in a Chalk Sandbox and stores them alongside the engineered features so every experiment starts from the same ready-to-train dataset.

Grid cut the loop from data preparation to evaluation by more than half and can compare models without introducing differences in data preparation.

Give every experiment the compute it needs, for training and evaluation.

Grid's data and ML team can launch training jobs themselves. A Grid ML engineer specifies the compute required for the run, and Chalk provisions compute when it starts and shuts it down when it ends. Every run logs the same metrics, plots and checkpoints, and the training run UI shows GPU utilization.

Grid now runs different models and hyperparameter settings as parallel jobs. Within days, the team trained Transformer and XGBoost candidates on the same platform and compared them side by side in a Chalk Notebook. There was no waiting period, and nothing to export or reconcile before deciding which model looked stronger. After moving the model development cycle to the Chalk platform, Grid no longer needed a separate MLflow setup.

“With Chalk Model Training, I can easily start a training run, and let Chalk handle the infrastructure. We can run different models and hyperparameters in parallel, which has allowed us to roughly triple the number of training runs we can complete.”
hi
Hyojin Lim Machine Learning Engineer

Dry-run the winning model, and then serve the new model in production.

Once Grid selects a winning model, one script turns the finished run into a versioned, serving-ready model. The model then scores live applicants in dry-run mode without changing real approvals. Grid can see how it ranks real traffic, and what that would mean for approvals and losses, before any decision depends on it.

Then, Grid serves the model with Chalk, and the new model makes real-time approval decisions using the same Chalk features it was trained on. With features defined once, Grid uses them across both training and real-time inference.

Now, a better model reaches production faster. For Grid, model development speed turns into business value: the faster the team deploys a better model, the sooner it can approve more qualified members at the same loss rate.

Close the loop with continuous retraining.

For their live models, Grid runs retraining as a scheduled Chalk Function that triggers on a monthly schedule. Each month, the job:

  1. Builds fresh training data from Chalk features. The function pulls the most recent feature data covering a rolling window of recent advances and adds a new version to Grid's training datasets. Dataset versions are tracked so any training result can be traced and reproduced.
  2. Trains a new candidate model. The function starts a Chalk Model Training run on managed compute infrastructure. The run trains the model, calibrates confidence scores, and saves a checkpoint.
  3. Evaluates the candidate against production. The current production model is scored on the same held-out rows as the new candidate for a direct comparison.
  4. Registers and reports. Each new version is recorded in the Chalk Model Registry. A Chalk log monitor posts the full metrics scorecard to Grid's Slack, comparing the new version with the production model.

A Grid engineer reviews the challenger and promotes it with one command. This process creates a recursive model development loop that runs independently, and engineers only need to step in to decide which new models to promote.

“Chalk has transformed how quickly we build and ship models. With scheduled training jobs and a continuous experimentation loop across training, evaluation, and serving, all on one platform, we’re developing models over 50% faster and taking ideas to production in days, not weeks.”
hi
Hyojin Lim Machine Learning Engineer

Chalk enables faster iteration and better underwriting decisions

Chalk cut Grid's end-to-end model development time by more than half and tripled its training velocity. The bigger impact is what the speed enables. Grid can test more ideas, identify better models faster, and put those models in front of real traffic sooner.

Eliminating infrastructure work means more time spent on the highest-value work: improving the models that make approval decisions.

“For our team, removing infrastructure overhead makes a meaningful difference. By eliminating the model training process overhead, parallelizing training runs, and creating a continuous improvement loop, Chalk helps us implement models that make better decisions faster. That ultimately gives us more confidence in how we scale the business.”
hi
Bijan Safai Software Engineering Manager

Now, experiments start when the idea is ready. Grid pays for compute only while a job runs. And the team can put two models head to head and trust the result, so the better model can be promoted to start approving more new Grid members sooner.

Learn more about the Chalk platform

With features, training, evaluation, and serving on one platform, Grid can take a model from notebook to live traffic without manually configuring infrastructure or bringing in extra help. The process becomes a continuous learning cycle, all on Chalk.

To get started, reach out to the Chalk team about your use case, and read the docs on Chalk Model Training.

Build faster with Chalk
See what Chalk can do for your team.