
Remote opportunity at
ChimeSenior Software Engineer, Machine Learning Platform
Chime, a financial technology company focused on providing helpful, easy, and free core financial services, is hiring a Senior Software Engineer, Machine Learning Platform . This…
Career Tools
About This Role
Chime, a financial technology company focused on providing helpful, easy, and free core financial services, is hiring a Senior Software Engineer, Machine Learning Platform . This fully remote position sits within the Machine Learning Platform team, which builds and operates the infrastructure, tooling, and developer experience powering machine learning throughout the enterprise. The role focuses on creating secure, reliable, and reusable platform capabilities for both traditional machine learning and emerging…
Job Description
Chime, a financial technology company focused on providing helpful, easy, and free core financial services, is hiring a Senior Software Engineer, Machine Learning Platform. This fully remote position sits within the Machine Learning Platform team, which builds and operates the infrastructure, tooling, and developer experience powering machine learning throughout the enterprise. The role focuses on creating secure, reliable, and reusable platform capabilities for both traditional machine learning and emerging AI workloads.
This opportunity is ideal for engineers experienced in distributed systems, cloud infrastructure, applied machine learning, and AI product engineering. You will collaborate closely with data science and machine learning engineering groups to design and operate scalable infrastructure on AWS, support LLM and agentic workloads, and shape technical strategy. Candidates must be comfortable participating in production on-call rotations and contributing to a highly collaborative, entrepreneurial environment that values integrity, diversity, and an owner's mindset.
Responsibilities
- Design, build, and operate scalable ML and AI infrastructure on Amazon Web Services
- Develop shared platform capabilities for large language model and agentic workloads encompassing model access, retrieval, state management, and orchestration
- Construct evaluation frameworks for non-deterministic AI systems including offline benchmarks, regression tests, and human feedback loops
- Establish reliability, governance, and observability standards covering traces, versioning, latency, safety, and cost efficiency
- Build distributed training, batch processing, and large-scale systems utilizing frameworks like Ray or Spark
- Manage infrastructure as code using Terraform and maintain feature stores and data pipelines
- Participate in on-call rotations to maintain production systems
- Partner with data science teams to improve developer experience and optimize continuous integration and deployment pipelines
Requirements
- Five or more years of professional experience in platform engineering, distributed systems, production machine learning systems, or AI infrastructure
- Familiarity with the complete machine learning development lifecycle from data preparation to monitoring
- Experience designing distributed systems and large-scale data or compute platforms on AWS using Spark or Ray
- Working knowledge of LLM application patterns including retrieval-augmented generation, tool calling, and agent orchestration
- Proficiency in infrastructure as code, DevOps practices, and CI/CD pipelines
- Containerization and orchestration expertise using Docker and Kubernetes
- Strong coding abilities in languages such as Python, Go, Java, or Scala
- Solid grounding in software engineering fundamentals including version control, testing, and code review
Qualifications
- Prior experience shipping production-grade LLM-powered or agentic systems
- Familiarity with model gateways, vector search, prompt lifecycle management, and tool execution frameworks
- Experience building tracing, evaluation, and observability tools for non-deterministic AI applications
- Knowledge of managed or self-hosted foundation model platforms such as Amazon Bedrock or SageMaker
- Background operating GPU-based workloads and optimizing training or inference performance and costs
Core Skills
Benefits
- Comprehensive health, financial, and wellbeing benefits
- Generous vacation policy alongside company-wide paid days off
- Up to 22 weeks of paid parental leave for birthing parents and 12 weeks for non-birthing parents
- Backup child, elder, and pet care along with subsidized commuter benefits
- Annual wellness stipend for eligible health-related expenses
- Access to family planning reimbursement and charitable time off programs
Frequently Asked Questions
What is the location and remote status for this role?
This is a fully remote position.
What is the salary range for the Senior Software Engineer, Machine Learning Platform position?
The base salary offered for this role ranges from $187,000 to $259,000 USD, with potential for bonuses, equity, and benefits.
What level of experience is required?
Candidates must have five or more years of experience in platform engineering, distributed systems, production ML systems, or AI infrastructure.
How can I apply if I need a disability accommodation?
If you require accommodation during the application process due to a disability or special need, you can contact accommodations@chime.com.
Sample Interview Questions
AI-generated questions tailored to this specific role — a preview of the full practice set.