Senior Software Engineer, OpenClaw Agent Platform
Build and operate the platform, runtime environments, and evaluation and observability systems powering Commure’s fleet of autonomous AI agents. The role requires strong Python, Linux, cloud, Kubernetes, and production-reliability experience, plus hands-on experience with LLM-powered agents.
About the job
Responsibilities
- Own the infrastructure and runtime environment for Commure’s fleet of OpenClaw agents.
- Build, configure, deploy, and troubleshoot agents running across Linux-based virtual machines.
- Develop Python services, automation, operational tooling, and agent integrations.
- Design, build, and operate agent harnesses supporting tool use, memory, context management, permissions, scheduling, retries, and human escalation.
- Create a standardized platform for engineers to develop, test, deploy, and update agents safely.
- Deploy and operate workloads across Linux virtual machines, Kubernetes, and Google Cloud or AWS infrastructure.
- Use Bash and Linux tooling to investigate system behavior, automate environment management, and resolve production issues.
- Build systems for scheduling, background execution, queues, concurrency management, and long-running tasks.
- Develop integrations with internal services, databases, APIs, communication platforms, and operational tools.
- Establish secure approaches to secrets, credentials, identity, permissions, and access control for autonomous agents.
- Build evaluation frameworks measuring reliability, task completion, tool accuracy, latency, and cost.
- Implement observability across agent traces, prompts, tool calls, infrastructure, failures, and resource consumption.
- Create dashboards, alerting, and incident-response processes.
- Diagnose failures across application code, model behavior, third-party tools, networks, containers, virtual machines, and cloud infrastructure.
- Improve platform efficiency and resilience as usage and workload complexity grow.
- Convert agent prototypes into dependable production systems.
- Establish best practices for developing and operating autonomous agents.
Requirements
- Professional experience building and operating production software, infrastructure, or internal platforms.
- Strong proficiency in Python.
- Excellent Linux experience, including development, deployment, and debugging on virtual machines.
- Strong Bash and command-line skills, including shell scripting, process management, networking diagnostics, filesystem operations, package management, and log investigation.
- Experience managing software across virtual-machine fleets.
- Production experience with Google Cloud Platform or Amazon Web Services, including compute, networking, storage, identity and access management, and observability.
- Experience deploying Linux virtual machines and containerized workloads in cloud environments.
- Hands-on experience with Kubernetes, Docker, and production container workloads.
- Experience building infrastructure, deployment systems, internal developer platforms, or distributed application runtimes.
- Experience with CI/CD, infrastructure as code, configuration management, secrets management, and automated deployments.
- Hands-on experience with large language models, autonomous agents, or agent harnesses.
- Understanding of tool calling, memory, context management, planning, retries, evaluations, permissions, and human-in-the-loop execution.
- Experience monitoring and debugging distributed systems with logs, metrics, traces, dashboards, and alerts.
- Understanding of production reliability, failure recovery, idempotency, rate limiting, concurrency management, and incident response.
- Ability to balance experimentation with production security and reliability.
- Strong ownership and communication skills across product, infrastructure, security, and operations teams.
Nice to Have
- Experience building, extending, deploying, or operating an OpenClaw agent.
- Contributions to OpenClaw or another open-source agent framework.
- Experience operating autonomous or long-running AI-agent fleets.
- Experience with multi-tenant agent platforms or internal developer platforms.
- Experience with agent evaluation, tracing, replay, simulation, or regression testing.
- Experience with browser agents, computer-use agents, or agents interacting with external systems.
- Familiarity with model gateways, inference providers, prompt versioning, token accounting, and LLM cost optimization.
- Experience securing workloads that access sensitive data or high-impact production tools.
- Familiarity with Swift and native Apple client applications.
- Experience in healthcare, revenue cycle management, or another regulated industry.
Skills
Python, Linux, Bash, Kubernetes, Docker, GCP, Amazon Web Services, CI/CD, Infrastructure As Code, Configuration Management, Secrets Management, LLMs, Autonomous Agents, Distributed Systems, Swift
Similar jobs
ML Engineering jobsBuild and operate production AI agents for finance workflows, owning orchestration, evaluation, reliability, and human-in-the-loop safeguards. Requires 8+ years building SaaS products, 2+ years shipping agentic systems, and hands-on Node.js and TypeScript experience.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.