DAT Solutions

Principal Software Engineer, ML Platform

Seattle, Washington, United States · Posted 2h ago

$257-321KprincipalpermanentDept: Tech - Data Platform (DAT610)

Job Description

About DAT

DAT Freight & Analytics is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five decades. Founded in 1978, DAT operates the largest freight marketplace in North America — processing 250 million+ load posts annually and maintaining one of the largest repositories of freight market transaction data in the world. On a defined path to $1 billion in revenue, DAT deploys a suite of software solutions, machine learning models, and intelligent automation tools that help brokers, carriers, and shippers price freight accurately, source capacity, reduce risk, and operate more efficiently. With nearly 700 teammates across offices in Denver, CO; Portland, OR; Seattle, WA; Springfield, MO; Toronto, ON; and Bangalore, India, DAT combines the credibility of a multi-decade market leader with the drive of a company that is not done disrupting the industry it helped build. For more information, visit www.DAT.com

 


Job Application Deadline: 09/30/2026

 

The Opportunity

DAT’s Science organization is seeking a Principal Software Engineer, ML Platform to lead the evolution of DAT’s most critical ML Platform capabilities.

As the platform enters a new phase of growth, we must increase our ability to experiment, learn, and adapt in real time across our marketplace, fraud detection, pricing, and other decision systems. We also need to scale and adapt ML capabilities built for sub-brands such as Convoy so they operate reliably within DAT’s broader product, data, and operational environment.

This role is both deeply hands-on and highly architectural. The Principal Engineer will lead a 3-4 member platform engineering team and set the technical direction for the foundational infrastructure that enables our ML and AI systems to iterate faster, adapt in real time, and operate safely at scale.

You will lead the development of the core capabilities that let us:

  • Deliver lower-latency data to models, unlocking online learning, adaptive policies, and improved real-time decision-making for our auction mechanisms, fraud detection systems, pricing workflows, and carrier engagement campaigns.
  • Evolve our ML platform to support generative AI, including orchestration, retrieval, standardized service patterns, and scalable model serving needed for foundational model applications in document digitization and voice-based features.
  • Experiment faster and safer through robust causal-inference tooling, richer randomized experimentation, and reliable evaluation infrastructure that helps us learn more about the unique spatio-temporal dynamics of a trucking marketplace.
  • Scale and operationalize ML models and platform capabilities for DAT’s scale, ensuring that differences in data, traffic, latency, reliability, and product integration are addressed systematically.

You will define and implement durable service architectures, build the real-time systems that power ML in production, lead the platform engineering team, and partner closely with scientists and product engineers to accelerate iteration and innovation.

Your work will form the backbone of the next generation of ML and AI capabilities across the unified freight network DAT represents.

What You’ll Do

As a Principal Software Engineer, ML Platform, you will lead the technical direction, execution, and evolution of the ML platform across DAT. You will manage the technical priorities of a 3-4 member platform engineering team, mentor engineers and scientists, and deliver solutions whose impact scales across teams and the broader organization, not just within individual projects.

Your work will influence three major areas:

Experimentation and Evaluation Infrastructure

Drive the evolution of DAT’s experimentation and model-evaluation foundations. Enable rigorous causal measurement, reliable online experimentation, scalable model iteration, and adaptive learning systems that continuously improve marketplace and policy decisions.

  • Evolve our various experimentation platforms to support richer randomized experiments, robust causal-inference tooling, and high-quality exposure and assignment logging; evaluate and integrate third-party solutions where beneficial.
  • Enable adaptive learning workflows, including reinforcement learning, contextual bandits, and online learning, for dynamic marketplace and policy decisions so scientists can safely run explore-exploit strategies in production.
  • Harden evaluation infrastructure across offline and online paths, including drift detection, guardrail metrics, and structured feedback loops that ensure reliable model behavior over time.
  • Develop orchestration patterns and tooling that combine inference, retrieval, business logic, policy guardrails, and human-in-the-loop steps into reusable, observable, and auditable workflows for ML-powered systems.
  • Establish cross-team standards for experimentation, evaluation, model feedback, and production readiness.

Feature Stores and Streaming Infrastructure

Lead the evolution of the low-latency feature store and real-time streaming platform that supports marketplace optimization, fraud detection, pricing, and other decision systems.

  • Iterate on and expand on existing low-latency feature store and real-time streaming platform (RisingWave) to deliver signals such as app analytics, carrier behavior, and digital fingerprints.
  • Ensure unified online and offline semantics to improve online decision-making, support real-time optimization, and enable future reinforcement-learning and online-learning workflows.
  • Build high-throughput streaming pipelines for carrier engagement, risk indicators, and fraud signals that power sub-minute marketplace and policy decisions.
  • Develop platform-level trucking knowledge systems, including RAG indexes, domain adapters, structured benchmarks, and retrieval strategies that ground AI systems in operational realities.
  • Lead the scaling and adaptation of ML capabilities for DAT’s traffic, data, reliability, and integration requirements.

Core DevOps and MLOps Foundations

Own the evolution of DAT’s ML Platform end to end, from data capture and transmission to storage and consumption by ML models and analytics. Reduce latency, increase reliability, and improve developer efficiency across teams.

  • Set the technical roadmap for the platform ecosystem, leveraging Kafka, Snowflake, Kubernetes, and modern data formats such as Avro, JSON, and Iceberg. Use Python and Go to build the connective tissue that ensures platform reliability and scale.
  • Build low-latency, production-grade services in Python and contribute to TypeScript/Node where needed, including emitting high-quality data signals, wiring model calls into product flows, and enabling experimentation and feature-flag pathways.
  • Partner with scientists and product engineers to define durable service patterns for API design, serving workflows, deployment, monitoring, and production ownership.
  • Mature core platform infrastructure, including Terraform/IaC, CI/CD, observability, logging and tracing, incident readiness, and cost/performance optimization.
  • Improve SQL/dbt workflows and batch and streaming pipelines to increase reliability, correctness, and scalability.
  • Extend model-serving infrastructure to support advanced ML workloads, from managed inference to self-hosted GPU, with standardized versioning, canary and A/B rollouts, and granular monitoring.
  • Establish engineering practices for safely scaling models and services from sub-brand environments into DAT production environments.

Leadership and Scope

  • Lead and mentor a 3-4 member platform engineering team, setting priorities and raising the quality bar for platform design, implementation, and operations.
  • Define and communicate the multi-year technical direction for the ML platform in partnership with Science, Product, and Engineering leadership.
  • Make architecture decisions that affect multiple teams and product surfaces across DAT
  • Identify and retire platform, reliability, data-quality, and operational risks before they constrain product delivery.
  • Build alignment across teams without relying solely on direct reporting authority.
  • Balance near-term product commitments with durable platform investments and reusable standards.

The Skills and Experience You’ll Bring

  • Extensive experience in ML engineering, data infrastructure, platform engineering, or closely related production engineering roles, with demonstrated impact at Principal or equivalent scope.
  • Deep hands-on experience with real-time ML platforms, including feature stores, stream processing, low-latency data services, and online inference systems.
  • Strong proficiency in Python, with the ability to work across non-Python stacks including TypeScript/Node, gRPC services, and Kubernetes-based microservice ecosystems.
  • Expertise in modern data and ML infrastructure, including Kafka, Kubernetes, Postgres-like OLTP systems, cloud platforms, and production observability tooling.
  • Experience building and operating robust data and ML pipelines, both batch and streaming, in high-scale environments such as marketplaces, fraud detection systems, pricing, personalization, or real-time decision platforms.
  • Strong DevOps and MLOps fundamentals, including CI/CD, containerization, infrastructure-as-code such as Terraform and Helm, automated monitoring, and cloud cost and performance optimization.
  • Track record of leading and mentoring a small engineering team while remaining hands-on in architecture and implementation.
  • Experience scaling and adapting production systems across businesses, platforms, or materially different operating environments.
  • Collaborative platform mindset, with a track record of partnering with scientists and product engineers to co-design durable service patterns for model serving, deployment, monitoring, and API design.
  • Ability to set technical direction across teams, identify and retire platform risk, and deliver solutions whose impact scales across the organization.

We'd Be Extra Excited If 

  • You have experience building ML systems in two-sided marketplaces, financial markets, or other economically complex environments, with intuition for incentives, pricing, and market dynamics.
  • You have deep experience with data reliability and correctness at scale, including schema evolution, data-quality enforcement, backfills, late-data handling, and incident response for production data systems.
  • You have applied advanced ML techniques such as reinforcement learning, bandits, or optimization to unlock real-world business impact, ideally within freight, logistics, or transportation technology.
  • You have led the integration or migration of ML platforms across brands, business units, or significantly different production environments.
  • You have experience building strong developer experience and platform adoption across multiple science and engineering teams.

 

Why DAT?
DAT is an award winning employer of choice.

For starters, we have a hybrid work environment, but we also know what makes a great workplace. We have a time-tested and resolute set of operating values predicated on integrity, mutual respect, open communication, and executing with excellence. These values inform our strategic vision as much as any one of our products does. We’ve been an employer of choice in the Portland metropolitan area for four decades, and within one year of opening our Denver office, DAT was #26 on Built In Colorado’s 100 Best Places to Work In Colorado.

  • Medical, Dental, Vision, Life, and AD&D insurance
  • Parental Leave
  • Flexible Vacation Time (FVT)
  • An additional 10 holidays of paid time off per calendar year
  • 401k matching (immediately vested)
  • Employee Stock Purchase Plan
  • Short- and Long-term disability sick leave
  • Flexible Spending Accounts
  • Health Savings Accounts
  • Employee Assistance Program
  • Additional programs - Employee Referral, Internal Recognition, and Wellness
  • Free TriMet transit pass (Beaverton Office)
  • Competitive salary and benefits package
  • Work on impactful projects in a cutting-edge environment
  • Collaborative and supportive team culture
  • Opportunity to make a real difference in the trucking industry
  • Employee Resource Groups

 

*This position is not eligible for visa sponsorship**

For Washington-based candidates, in compliance with the Washington State Pay Transparency Law, the salary range for this role is $257,000.00 - $321,400.00 + target bonus.  DAT considers factors such as scope and responsibilities of the position, candidate's work experience, education and training, core skills, internal equity, and market and business elements when extending an offer.

 

DAT embraces the value of a diverse workforce, and believes it is a core strength of our company that we encourage those values in every DAT employee, at every level of our organization, regardless of tenure or rank. We provide equal employment opportunities (EEO) to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, amnesty, or status as a covered veteran in accordance with applicable federal, state, and local laws.

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities

The contractor will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by the employer, or (c) consistent with the contractor’s legal duty to furnish information. 41 CFR 60-1.35(c)

#LI-RF1

#LI-hybrid