Logo for iCIMS

Sr. Principal Engineer, SRE

Role overview

Qualifications

  • Bachelor’s degree in computer science, Engineering, Information Systems, or related technical field
  • 8+ years in SRE, DevOps, Infrastructure Engineering, or Observability Engineering roles with 4+ years in senior technical positions
  • Proven hands-on experience designing, implementing, and operating observability capabilities at scale
  • Strong multi-cloud and cloud-native experience, including AWS, containers, Kubernetes/ECS, Linux, and distributed application architectures

Responsibilities

  • Provide strategic technical direction for a team of 5+ SRE engineers
  • Participate in enterprise-wide incident management, ensuring rapid detection and response
  • Establish and evolve enterprise standards for logs, metrics, traces, and service ownership
  • Design and operate scalable observability platforms and telemetry pipelines

About the company

iCIMS logo

iCIMS

Computer Software / SaaS

Ideagen plc provides market-leading information management, safety, risk and compliance software solutions that allow organisations to achieve operational excellence, regulatory compliance and reduce risk. The Group has shown excellent growth, both organically and through strategic acquisitions, and is listed on The London Stock Exchange AIM market (Ticker: IDEA.LN). As authors of an excellent portfolio of software products, the Group is able to provide complete content lifecycle solutions that enable organisations to meet their Regulatory and Quality Compliance standards, helping them to reduce costs and improve efficiency. Our Mission Statement is: “To enable our clients to improve their organisations by providing the tools which can help improve customer service, increase efficiency, reduce risk, enhance compliance, and lower costs" Ideagen's wide portfolio of solutions range from Audit & Risk Management, Document Management & Workflow, Capture, Process Mapping, Order Communications, Infection Control, Electronic Medical Record and ED Management. For more information, visit our website >> www.ideagen.com or contact us via email info@ideagen.com or telephone +44 (0) 1629 699100.

Company details

Company typeLarge
IndustryComputer Software / SaaS
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Overview:

We are seeking an experienced Sr. Principal Engineer, Site Reliability (SRE) to establish and scale the observability and reliability foundation for our global multi-cloud SaaS platform serving thousands of customers worldwide. This role will define enterprise standards for logging, metrics, tracing, alerting, dashboards, SLIs/SLOs, and service ownership, then work hands-on with Engineering and Operations teams to put those standards into practice. The successful candidate will bring deep experience operating observability at scale, modernizing telemetry and tooling, reducing alert noise and tool sprawl, and using actionable signals to improve detection, MTTR, performance, and reliability. Off hours support as needed

 

Success Metrics

  • Customer Impact: Reduced MTTD/MTTR and improved customer experience through faster detection, diagnosis, and recovery
  • Observability Adoption: Measurable adoption of common logging, metrics, tracing, alerting, dashboarding, and service ownership standards across critical services
  • Reliability Engineering: Expanded use of meaningful SLIs, SLOs, and error budgets to drive service health and engineering priorities
  • Signal Quality: Reduction in noisy, duplicate, and non-actionable alerts while improving coverage of critical customer journeys and dependencies
  • Tooling & Cost Efficiency: Improved observability platform efficiency through governance, consolidation, telemetry optimization, and reduced tool sprawl
  • Cross-functional Adoption: Strong partnership with Product and Engineering teams that translates standards into measurable production adoption
About Us:

ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage. For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences. 

 

ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact. 

Responsibilities:

Technical Leadership

  • Provide strategic technical direction for a team of 5+ SRE engineers across one or more geographic regions (US, Ireland, or India)
  • Own the technical strategy and roadmap for enterprise observability and reliability capabilities in partnership with SRE, Engineering, Cloud, and Product teams
  • Define reference architectures, engineering patterns, and standards that teams can consistently apply in production
  • Drive architecture reviews and technical decision-making for complex observability, reliability, scalability, and performance challenges
  • Provide hands-on technical mentorship and guidance, raising observability and SRE engineering practices across teams

Incident Management & Response

  • Participate in enterprise-wide incident management, ensuring rapid detection, response, restoration, and prevention of recurring issues
  • Improve incident detection and triage through actionable telemetry, service health views, dependency context, and well-designed alerting
  • Develop and maintain runbooks, emergency response procedures, and operational readiness practices for critical services
  • Lead root cause and post-incident reviews, ensuring clear documentation and implementation of durable corrective actions
  • Participate in 24/7 on-call and escalation procedures and serve as a senior technical leader with Engineering and Incident Management during critical incidents

Observability Strategy & Standards

  • Establish and evolve enterprise standards for logs, metrics, traces, alerting, dashboards, instrumentation, and service ownership
  • Champion OpenTelemetry-first, vendor-neutral instrumentation patterns with consistent context, correlation, naming, and metadata across services
  • Implement meaningful SLIs, SLOs, error budgets, and service health views that connect technical signals to customer impact
  • Drive practical adoption and governance of observability standards, measuring coverage, signal quality, and operational effectiveness across teams

Platform Reliability, Automation & Tooling

  • Design and operate scalable observability platforms and telemetry pipelines using technologies such as Grafana, Prometheus, Sumo Logic, New Relic, and cloud-native services
  • Lead observability platform modernization, migration, and consolidation while maintaining coverage and controlling ingestion, retention, cardinality, and overall tooling cost
  • Use infrastructure-as-code, automation, self-service patterns, and automated remediation to make reliability practices repeatable and reduce operational overhead
  • Monitor and optimize multi-cloud infrastructure and core services across AWS, Azure, and GCP for reliability, performance, capacity, and operational efficiency
Qualifications:
  • Bachelor’s degree in computer science, Engineering, Information Systems, or related technical field
  • Equivalent combination of education and experience will be considered
  • Cloud certifications (AWS, Azure, or Google Cloud)

Technical Experience

    • 8+ years in SRE, DevOps, Infrastructure Engineering, or Observability Engineering roles with 4+ years in senior technical positions
    • Proven hands-on experience designing, implementing, and operating observability capabilities at scale in large enterprise SaaS or cloud production environments
    • Deep experience across logging, metrics, distributed tracing, alerting, dashboards, and OpenTelemetry, with platforms such as Grafana, Prometheus, Sumo Logic, and New Relic
    • Strong multi-cloud and cloud-native experience, including AWS, containers, Kubernetes/ECS, Linux, and distributed application architectures
    • Experience designing scalable telemetry pipelines and managing sampling, retention, cardinality, data quality, and cost tradeoffs in high-volume environments

Leadership & Communication

    • Proven track record creating technical standards and successfully driving them from architecture into consistent production adoption across engineering teams
    • Experience serving as a senior technical leader during critical incidents and complex cross-team reliability initiatives
    • Strong communication and influencing skills with engineers, architects, product leaders, and senior stakeholders
    • Demonstrated ability to mentor technical teams, build alignment across organizational boundaries, and lead through influence

SRE & Operations

    • Demonstrated success implementing SRE principles in large-scale production environments, including practical use of SLIs, SLOs, and error budgets
    • Strong background in incident management, root cause analysis, operational readiness, automation, and continuous reliability improvement
    • Experience with ITIL frameworks and tools and with establishing service-level expectations for enterprise SaaS products
Preferred:
    • Experience leading large-scale observability platform migrations, consolidation initiatives, and telemetry cost/governance programs
    • Infrastructure-as-code expertise with Terraform or CloudFormation; authentication and identity management systems knowledge is a plus
EEO Statement:

iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at iCIMS. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you’d like to request an accommodation due to a disability, please contact us at careers@icims.com.

Compensation and Benefits:

Competitive health and wellness benefits include medical insurance (employee and dependent family members), personal accident and group term life insurance, bonding and parental leave, lifestyle spending account reimbursements, wellness services offerings, sick and casual/emergency days, paid holidays, tuition reimbursement, retirals (PF - employer contribution) and gratuity. Benefits and eligibility may vary by location, role, and tenure. Learn more here: https://careers.icims.com/benefits

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at iCIMS

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.