Software Engineer III, Agentic Systems and Production Safety
- New York City, Kirkland, United States
- $147,000 – $210,000
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 2 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
- Experience in Backend software development and Data Analysis.
- Experience in C++, Java or Go.
Preferred qualifications:
- Master's degree or PhD in Computer Science or related technical fields.
- 2 years of experience with performance, large-scale systems data analysis, visualization tools, or debugging.
- 2 years of experience with data structures and algorithms in either an academic or industry setting.
- Proficiency in code and system health, diagnosis and resolution, and software test engineering.
- Understanding of the Generative AI/Large Language Model (LLM) landscape, including model evaluations, model capabilities, prompt engineering, fine-tuning, and agentic architectures.
About the job
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.
The Access Transparency (AXT) team is foundational to customer trust in Google Cloud Platform (GCP), ensuring every administrative access to user content is justified, transparent, and compliant with critical global standards. Within AXT, our team operates at the intersection of AI compliance and site reliability. We lead initiatives to eliminate operational risks by building self-healing automation and advanced telemetry that prevent unsafe manual interventions in production. We also build the foundational infrastructure that ensures autonomous AI workflows and agents remain verifiable, transparent, and compliant with strict regulatory standards. Joining this team means handling high-impact, distributed systems challenges that define how both human engineers and AI agents safely operate across Google Cloud.
Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $147000 - $210000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities
- Write product or system development code.
- Build industry-first advanced logging and distributed tracing systems that ensure autonomous workflows and AI agents are transparent and compliant with strict regulatory standards.
- Design scalable technical architectures and produce well-tested code (primarily in Go, C++, or Java). Leverage AI-powered coding assistants to accelerate routine development, empowering you to focus your bandwidth on high-level system design and complex problem-solving.
- Design data pipelines, monitoring tools, and enforcement systems that enforce least-privilege access and significantly reduce the risk of manual operations in production environments.
- Partner with Site Reliability Engineers (SREs) and developers across the organization to replace operational toil with self-healing automation.
Skills
- C++
- Java
- Go
- Distributed Systems
- Large-Scale System Design
- Data Analysis
- Artificial Intelligence





