Infrastructure Monitoring Engineer
Job description
About the role
This position operates within the Network Operations Center (NOC) at PubMatic, serving as a critical line of defense for the company's digital advertising infrastructure. The hire will be responsible for sustaining high availability, performance, and reliability for all production systems that facilitate real-time ad transactions. You will leverage advanced AI-driven tools and automation frameworks to strengthen incident detection capabilities and improve troubleshooting efficiency across complex environments. A core part of this role involves using modern AI assistants to accelerate root cause analysis, deep log examination, and the generation of actionable operational insights. You will transform raw data into clear narratives of system behavior, ensuring that every shift contributes to a more resilient infrastructure. The position demands a proactive mindset, focusing on identifying issues before they impact the advertising ecosystem. You will act as a bridge between technical systems and business outcomes, ensuring that uptime directly supports revenue flow. This role is fundamental to the operational excellence of PubMatic's sell-side platform.
Key facts
What you'll do
You will monitor infrastructure, applications, and network systems using dashboards and platforms such as Grafana and Nagios to maintain situational awareness. You will handle alerts and incidents classified as P1 and P2, performing initial triage, diagnosis, and ensuring timely escalation and resolution to minimize business impact. You will provide Tier-1 support for production systems and services while maintaining precise shift handovers to ensure continuity and context preservation. You will collaborate with Engineering, Ad Operations, and DevOps teams to isolate faults, correlate dependencies, and resolve issues efficiently. You will support real-time systems involved in ad serving, bidding, and traffic flow to ensure that user demand is met without interruption. You will utilize AI tools such as log analysis assistants and alert summarization instruments to speed up debugging and root cause analysis, reducing mean time to resolution. You will craft and apply structured prompts to reliably extract technical answers and insights from logs, metrics, and incident data repositories. You will participate in deployment monitoring and post-release validation to verify that new changes do not introduce regressions or performance degradation. You will document incidents in detail, contribute to root cause analysis reports, and maintain updated operational runbooks and Wiki content for team knowledge sharing. You will analyze trends to identify recurring issues and suggest automation or AI-assisted solutions that improve system resilience and team efficiency.
Requirements
You must hold 1 to 3 years of hands-on experience in NOC, infrastructure monitoring, or production support roles within a technology-driven environment. You must possess a basic understanding of Linux, including command line operations, processes, memory management, disk I/O, and networking fundamentals. You must have current familiarity with monitoring tools such as Grafana and Nagios or similar platforms used for observability and alerting. You must have a working knowledge of networking concepts, including TCP/IP protocols, DNS resolution, and basic routing principles to troubleshoot connectivity issues. You must practice incident management methodologies and be familiar with alerting and ticketing systems such as Jira, Zenduty, or comparable workflow platforms. You must demonstrate the ability to use structured prompts and templates that reliably extract accurate technical answers from large language models. You must maintain a strong attention to detail to ensure that monitoring configurations, alert thresholds, and documentation are accurate and current.
Nice to have
Hands-on experience with AI tools such as ChatGPT for analysis, logging, summarization, and automated documentation tasks.
Skills & tools
Grafana, Nagios, Jira, Zenduty, log analysis assistants, and advanced prompt engineering techniques.
Practical notes
This role is based in Pune, India, and is an Infrastructure monitoring position within the NOC for real-time ad flow. The compensation offered is a base salary of 18 LPA. There are no specific travel requirements, visa sponsorships, or application deadlines mentioned in the source material. The engagement is focused on maintaining the stability and performance of PubMatic's advertising infrastructure through continuous monitoring and rapid incident response.