Logo for ICONMA

Lead SRE – Observability | Splunk | OpenTelemetry | Kubernetes

Role overview

Qualifications

  • 8+ years in Site Reliability Engineering, Platform Engineering, DevOps, or Observability
  • Experience designing and operating enterprise-scale observability platforms
  • Strong knowledge of Splunk SPL
  • Experience with Terraform and Infrastructure as Code

Responsibilities

  • Design, deploy, and operate enterprise observability platforms
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search
  • Develop dashboards, alerts, analytics, and trace visualizations

About the company

ICONMA logo

ICONMA

Management Consulting

We provide Professional Staffing Services & Project-Based Solutions for a broad range of Fortune 500 organizations. ICONMA is a certified Woman-Owned staffing company and was founded in 2000. ICONMA’s corporate headquarters is in Troy, Michigan, and has 15+ locations worldwide. What makes ICONMA stand out in a fiercely competitive industry? *We provide integrated, full lifecycle services across a broad range of business and technical platforms. *No single company can duplicate our full range of staffing and permanent recruiting services nationwide. *Proven track record of attracting and retaining exceedingly skilled professional workers in a highly competitive market. SERVICES OFFERED Staff Augmentation (Contract, Contract to Hire, Direct Hire, Single Source) Data Analysis Project-Based Services & Solutions Hire Train Deploy Service Model Offshore Staff Augmentation Payroll Services AREAS OF EXPERTISE - Information Technology - Engineering - Business Professional - Accounting/Finance - Admin/Clerical/Call Center - Healthcare/Clinical/Scientific - Marketing/Creative mail linkedin@iconma.com Phone (888) 451-2519 Website http://www.iconma.com

Company details

Company typeLarge
IndustryManagement Consulting
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Our client, a IT Services and Consulting company, is looking for a Lead SRE – Observability | Splunk | OpenTelemetry | Kubernetes for their Remote location.
 
Responsibilities:

  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.
 
Requirements:
  • 8+ years in Site Reliability Engineering, Platform Engineering, DevOps, or Observability.
  • Experience designing and operating enterprise-scale observability platforms.
Preferred Domain Background
  • Enterprise SaaS
  • Cloud Infrastructure
  • Networking
  • Cybersecurity
  • Large-scale Platform Engineering
  • FedRAMP or regulated cloud environments
  • Technology Stack :
  • Unix, Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
  • Experience supporting FedRAMP or regulated environments.
  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.
  • Years of Experience: 14.00 Years of Experience
Skills: 
  • Category          Name   Required          Importance       Experience
  • Cognitive Search          Elasticsearch    Yes      1                     
  • Data Management        Splunk Yes      1                     
  • DevOps           SRE     Yes      1                     
  • Iaas:Automation & Analytics    Grafana Yes      1                     
  • IDE Application Architect        DevOps           Yes      1                     
  • Infrastructure Services  Kubernetes       Yes      1                     
  • Quality Engineering      Ruby    Yes      1                     
  • Software Skills UNIX   Yes      1                     
 
Why Should You Apply?

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Other jobs at ICONMA

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.