Logo for Hard Rock Digital

Senior Site Reliability Engineer

Role overview

Qualifications

  • Deep infrastructure expertise
  • Experience with AI-driven operations
  • Expertise in Java-based applications
  • Knowledge of observability and monitoring tools

Responsibilities

  • Ensure the availability, reliability, and performance of high-traffic Java-based applications
  • Deploy and manage the Grafana stack for real-time monitoring and logging
  • Design and operate agentic AI workflows for operational tasks
  • Collaborate with cross-functional teams to improve reliability and leverage AI capabilities

About the company

Hard Rock Digital logo

Hard Rock Digital

Sports Betting & iGaming

Hard Rock Digital is building the future of online sports betting and interactive gaming. We’re turning up the volume on sports betting in a bold new way. Known the world over for our famous cafes, casinos, hotels, and rock memorabilia collection, our newest venture takes the same Hard Rock ethos and brings it to the newly expanding online sports betting industry here in the USA. Headquartered in Hollywood, Florida, with offices in New Jersey and Texas, we are a fast-paced team dedicated to building an unrivaled betting experience for millions of sports fans everywhere.

Company details

IndustrySports Betting & iGaming
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Location: Poland only, fully remote

Job Type: B2B, full time

 

Overview

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We care about each customer's interaction, experience, behaviour, and insight and strive to ensure we’re always acting authentically.

 

Rooted in the kindred spirits of the Seminole Tribe of Florida, the new Hard Rock Digital taps a brand known all over the world as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space.

 

What’s the position?

We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents.

 

You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs.

 

Key Responsibilities

Application Reliability & Performance

  • Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment.

  • Troubleshoot and resolve complex issues across production and non-production environments.

  • Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance.

  • Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling.

 

Monitoring, Observability & AIOps

  • Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting.

  • Implement and refine observability strategies that enhance visibility into application and infrastructure health.

  • Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring.

  • Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction.

 

AI & Agentic Workflow Engineering

  • Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization.

  • Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval.

  • Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents.

  • Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems.

  • Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving.

  • Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization.

 

Incident Management & Root Cause Analysis

  • Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence.

  • Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents.

  • Document and share lessons learned, contributing to a culture of continuous improvement.

 

Automation & Toil Reduction

  • Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements.

  • Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language.

  • Measure and report on toil reduction metrics to quantify the impact of automation initiatives.

 

Collaboration & Cross-functional Support

  • Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities.

  • Collaborate with DevOps and NOC teams to support the application platform.

  • Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders.

  • Provide feedback on application performance, potential improvements, and observability metrics.

 

Why This Role Is Different

This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Hard Rock Digital

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.