# Senior Software Engineer, OpenClaw Agent Platform

**Company:** [Commure](https://hotfix.jobs/companies/commure)
**Location:** Rio de Janeiro, Brazil
**Role:** ML Engineering
**Salary:** $50k – $93k/yr
**Experience:** 5+ years
**Skills:** Python, Linux, Bash, Kubernetes, Docker, GCP, Amazon Web Services, CI/CD, Infrastructure As Code, Configuration Management, Secrets Management, LLMs, Autonomous Agents, Distributed Systems, Swift
**Posted:** 2026-07-30

> Build and operate the platform, runtime environments, and evaluation and observability systems powering Commure’s fleet of autonomous AI agents. The role requires strong Python, Linux, cloud, Kubernetes, and production-reliability experience, plus hands-on experience with LLM-powered agents.

## Job Description

## Responsibilities
- Own the infrastructure and runtime environment for Commure’s fleet of OpenClaw agents.
- Build, configure, deploy, and troubleshoot agents running across Linux-based virtual machines.
- Develop Python services, automation, operational tooling, and agent integrations.
- Design, build, and operate agent harnesses supporting tool use, memory, context management, permissions, scheduling, retries, and human escalation.
- Create a standardized platform for engineers to develop, test, deploy, and update agents safely.
- Deploy and operate workloads across Linux virtual machines, Kubernetes, and Google Cloud or AWS infrastructure.
- Use Bash and Linux tooling to investigate system behavior, automate environment management, and resolve production issues.
- Build systems for scheduling, background execution, queues, concurrency management, and long-running tasks.
- Develop integrations with internal services, databases, APIs, communication platforms, and operational tools.
- Establish secure approaches to secrets, credentials, identity, permissions, and access control for autonomous agents.
- Build evaluation frameworks measuring reliability, task completion, tool accuracy, latency, and cost.
- Implement observability across agent traces, prompts, tool calls, infrastructure, failures, and resource consumption.
- Create dashboards, alerting, and incident-response processes.
- Diagnose failures across application code, model behavior, third-party tools, networks, containers, virtual machines, and cloud infrastructure.
- Improve platform efficiency and resilience as usage and workload complexity grow.
- Convert agent prototypes into dependable production systems.
- Establish best practices for developing and operating autonomous agents.

## Requirements
- Professional experience building and operating production software, infrastructure, or internal platforms.
- Strong proficiency in Python.
- Excellent Linux experience, including development, deployment, and debugging on virtual machines.
- Strong Bash and command-line skills, including shell scripting, process management, networking diagnostics, filesystem operations, package management, and log investigation.
- Experience managing software across virtual-machine fleets.
- Production experience with Google Cloud Platform or Amazon Web Services, including compute, networking, storage, identity and access management, and observability.
- Experience deploying Linux virtual machines and containerized workloads in cloud environments.
- Hands-on experience with Kubernetes, Docker, and production container workloads.
- Experience building infrastructure, deployment systems, internal developer platforms, or distributed application runtimes.
- Experience with CI/CD, infrastructure as code, configuration management, secrets management, and automated deployments.
- Hands-on experience with large language models, autonomous agents, or agent harnesses.
- Understanding of tool calling, memory, context management, planning, retries, evaluations, permissions, and human-in-the-loop execution.
- Experience monitoring and debugging distributed systems with logs, metrics, traces, dashboards, and alerts.
- Understanding of production reliability, failure recovery, idempotency, rate limiting, concurrency management, and incident response.
- Ability to balance experimentation with production security and reliability.
- Strong ownership and communication skills across product, infrastructure, security, and operations teams.

## Nice to Have
- Experience building, extending, deploying, or operating an OpenClaw agent.
- Contributions to OpenClaw or another open-source agent framework.
- Experience operating autonomous or long-running AI-agent fleets.
- Experience with multi-tenant agent platforms or internal developer platforms.
- Experience with agent evaluation, tracing, replay, simulation, or regression testing.
- Experience with browser agents, computer-use agents, or agents interacting with external systems.
- Familiarity with model gateways, inference providers, prompt versioning, token accounting, and LLM cost optimization.
- Experience securing workloads that access sensitive data or high-impact production tools.
- Familiarity with Swift and native Apple client applications.
- Experience in healthcare, revenue cycle management, or another regulated industry.

## Similar jobs

- [Senior Applied AI Software Engineer](https://hotfix.jobs/jobs/edd384fa-194d-4ebb-9c16-de26dbcf5680) - Kinter - Remote
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Agent Systems Engineer](https://hotfix.jobs/jobs/2b04a068-018b-4fe0-915f-4fdf80d4c582) - Adaption Labs - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/fe20f4a6-526c-475d-8c28-39b591dc52c4
**Canonical:** https://hotfix.jobs/jobs/fe20f4a6-526c-475d-8c28-39b591dc52c4