Senior Infrastructure Engineer
Senior Infrastructure Engineer owns critical infrastructure decisions, builds scalable platforms using AWS and Terraform, ensures security/compliance, and mentors teams. Requires 5+ years AWS experience and expertise in monitoring tools like Datadog.
About the job
What You’ll Do
- Design, build, and maintain core platform infrastructure to support scalable, reliable, and secure services across engineering teams.
- Develop and manage Infrastructure as Code (IaC) using tools like Terraform to ensure consistent, reliable environments.
- Own major architectural decisions and long-term technical direction for Fieldguide’s infrastructure.
- Build reusable platforms, abstractions, and AI-enabled tooling that raise the baseline for engineering teams.
- Monitor and improve system reliability, performance, and cost efficiency through metrics, logging, and alerting frameworks (e.g., Datadog, CloudWatch).
- Ensure infrastructure security and compliance by implementing best practices for identity management, network segmentation, secrets handling, and vulnerability management.
- Lead incident response and postmortem processes, driving root cause analysis and structural long-term improvements.
- Mentor and collaborate with engineers and tech leads, fostering a culture of reliability, automation, and continuous improvement.
- Support disaster recovery and business continuity planning, ensuring high availability and resilience of critical systems.
- Document infrastructure design, architecture decisions, and operational procedures for transparency and team enablement.
Who You Are
- You have 5+ years of hands-on experience constructing complex cloud solutions using multiple AWS services.
- You are skilled in provisioning and configuring cloud services using Terraform and the AWS CLI / API.
- You have proficiency in designing effective monitoring/alerting and log aggregation solutions using tools like Datadog and AWS CloudWatch (New Relic, Prometheus/Grafana, etc.)
- You have a solid understanding of data systems, including both SQL and NoSQL.
- You have experience in developing and maintaining software in security and regulatory compliance environments (SOC 2, PCI-DSS, HIPAA, etc.)
- You are comfortable participating in on-call support to ensure 24/7 availability of services.
- You have a passion for mentoring and coaching other engineers.
- You have excellent communication and organizational skills and are capable of managing multiple competing priorities.
- You have deep expertise designing systems and processes that make engineering teams measurably faster and more effective.
- You can clearly communicate technical strategy to managers and executives.
Bonus Points
- You have experience with GraphQL as a database front-end API.
- You have experience with database system architecture (e.g., Postgres) and observability, to help us increase our overall database performance.
- You have experience both working with AI, and with providing it as a tool for engineers and our internal applications to utilize.
- You have experience working through and designing for security audits (e.g., SOC2, PCI, etc.)
Skills
AWS, Terraform, Datadog, Aws Cloudwatch, Prometheus, Grafana, Kubernetes, SQL, NoSQL, Postgres
Similar jobs
DevOps / SRE jobsSenior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.
Own and optimize ultra-low-latency network infrastructure for institutional trading across cloud, on-premises, and colocated environments. The role requires 8+ years of network or infrastructure experience, deep routing and multicast expertise, production incident leadership, and automation skills.
Build and own Webflow’s corporate cloud foundation, including landing zones, networking, security, Infrastructure as Code, GitHub delivery pipelines, self-service deployment patterns, and observability. The role requires 5+ years of platform or cloud engineering experience and strong AWS, Azure, or GCP expertise.
Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.