Logo for Johnson Technology Systems Inc

Site Reliability Engineer

Role overview

Qualifications

  • Expert with Kubernetes
  • Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters
  • Strong understanding of distributed systems, search platforms, indexing pipelines
  • Excellent communication and prioritization skills

Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services
  • Administer and optimize OpenSearch environments for high availability and performance
  • Analyze and resolve operational issues, platform instability, and production incidents
  • Develop and maintain monitoring policies and operational runbooks

About the company

Johnson Technology Systems Inc logo

Johnson Technology Systems Inc

IT Services & IT Consulting

Driving Digital Transformation with Cutting-Edge IT & Engineering Solutions Established in 2003, JTSi (Johnson Technology Systems Inc.) is a trusted IT and Engineering Services provider with a strong track record of delivering mission-critical solutions to both the Public and Private Sectors. We specialize in Digital Transformation, Cloud Migration, SAP RISE implementation, S/4HANA upgrades, SAP BTP solutions (Concur, SuccessFactors), SAP AI (Joule) strategy, and Managed Services (MSP)—helping businesses and government agencies optimize operations and achieve success. With a customer-first approach, we don’t just offer services; we build long-term partnerships that drive innovation and efficiency. Our solutions are backed by an experienced team of cleared professionals, including PMI, ITIL & Black Belt-certified managers, ensuring secure, scalable, and high-performing IT environments. Why JTSi? ✔ Trusted by the U.S. Department of Defense since 2003 ✔ Proven success in Government & Commercial Sectors ✔ ISO 20000-1 & CMMI-L3 Certified | ITAR Registered ✔ Highly successful Cooperative R&D Agreement with the U.S. Army ✔ 97% rating in D&B/Open-Ratings Supplier Performance Evaluation ✔ Expertise in SAP, AI, Cloud, and Managed Services ✔ Cleared Facilities with Dedicated Security Officers Onsite ✔ University Innovation Center Partnerships From SAP RISE implementation to Cloud Migrations and AI-driven automation, JTSi is at the forefront of innovation. Our E2E cloud-based processes and MSP solutions enable organizations to scale efficiently while staying compliant with DoD, FedRAMP, CMMC, and ITAR standards. At JTSi, "Customer is Always First." We do what we say—ensuring reliability, security, and unmatched excellence in IT services. Explore the future with JTSi. Connect with us today!

Company details

IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

"WE DO WHAT WE SAY "

JTSi is a federal government consulting firm, providing technical services to the Federal Government, i.e., DoD, Client and various Civilian Agencies. We are proud to have earned the reputation of honesty, integrity and the ability to build long-term professional relationships with our employees and clients. Please visit our website at www.JTSUSA.com to learn more about who we are and what we do.

Company Name: - JTSi (Johnson Technology Systems, Inc.)
Title: Site Reliability Engineer
Location: Remote
Salary : $155K - $160K / year

DESCRIPTION OF PROJECT AND TASKS:

- MUST be a US Citizen & ONLY hold US Citizenship (No Dual Citizens)
- Fully remote position
- Possible convert to hire after 1 year
- Only submit candidates that have OpenSearch deployment ON KUBERNETES experience

Site Reliability Engineer – OpenSearch

Client is seeking a Site Reliability Engineer – OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services. This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment.

The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.

General Responsibilities

- Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environment
- Architect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground up
- Administer and optimize OpenSearch environments for high availability, resiliency, scalability, security, and performance
- Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization
- Analyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layers
- Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement
- Maintain the integrity and security of servers, systems, and OpenSearch platform infrastructure
- Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery
- Develop and maintain monitoring policies, alerting standards, operational runbooks, and support procedures
- Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services
- Ensure proper resource allocation and capacity planning across compute, memory, storage, and network resources
- Partner with product development and engineering teams to design and enhance service reliability and operational readiness
- Develop and implement testing strategies and document results for platform changes and operational improvements
- Support log ingestion, index management, retention policies, lifecycle management, and search performance tuning
- Work in a diverse environment and cross-train with other global team members
- Participate in an on-call rotation and support weekend or after-hours operational needs as required

Requirements
- Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.
- Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
- Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
- Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
- Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
- Expertise with Git
- Expertise with Concourse, including setup, management, and troubleshooting of new pipelines
- Expertise with Linux, specifically SUSE and Ubuntu
- Expertise with Kafka, Zookeeper, and Big Data technologies
- Expert in development of automation for testing, deployment, scalability, and management of cloud services
- Expertise with building, implementing, and/or supporting cloud monitoring tools
- Expert knowledge of cloud computing, infrastructure operations, and databases
- Expert understanding of web services, networking, virtualization, and internet protocols
- Ability to multitask and handle various projects, deadlines, and changing priorities
- Excellent communication and prioritization skills
- Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
- Strong interpersonal, presentation, and customer service skills

Desired Qualifications
- Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
- Experience deploying and operating OpenSearch in AWS-based environments
- Experience with Cloud Foundry-based environments
- Experience with Jenkins, Chef, and/or Terraform
- Exposure to and understanding of troubleshooting IP networks and application stacks
- Experience with observability tools such as Prometheus and Grafana
- Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
- Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms

Education
- BS/BA degree in Computer Science, Management Information Systems, or related IT discipline preferred
- Allowable substitution: An additional four (4) years of experience may be substituted for a BS/BA degree
- 8+ years of experience

Additional Requirements
- Participation in an on-call rotation for handling P1 incidents is required
- Flexible schedule which may include weekend or after-hours work
- Ability to work effectively in a diverse, collaborative, and globally distributed team environment


We recruit, employ, train, compensate and promote without regard to race, religion, color, citizenship, national origin, age, sex, gender, gender identity/expression, sexual orientation, marital status, disability, genetic information, veteran status or any other characteristic protected by federal, state, or local law.

Disclaimer: Nothing in this job description/posting shall constitute an offer or promise of employment. If you are not reviewing this job posting on our Careers' site http://jtsusa.com/careers or one of our approved job boards we cannot guarantee the validity of this posting. For a list of our current postings, please visit us at http://jtsusa.com/careers

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Johnson Technology Systems Inc

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.