Infrastructure Engineer (iGaming)
BETBY, Remote
As a fast-growing, award-winning company, Betby powers the industry with our premium sportsbook, featuring world-class risk management and seamless omni-channel support, reaching millions of players across countless markets.
Responsibilities:
— Designing, deploying, configuring, and maintaining scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments;
— Operating Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning;
— Building and maintaining reliable metrics and log collection pipelines for infrastructure and business-critical services;
— Creating dashboards that provide clear, useful visibility into service health, performance, capacity, and operational risks;
— Designing, tuning, and maintaining actionable alert rules and notification routing; reducing alert noise and improving incident response;
— Monitoring infrastructure and application metrics and logs, troubleshooting issues, and improving stability and performance under heavy loads;
— Managing metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective;
— Establishing high-availability and recovery approaches for observability services and validating operational readiness;
Requirements:
— Minimum 3 years of experience with administering Linux systems and operating monitoring, logging, or observability systems;
— Experience with Debian-based systems;
— Experience with Docker and Kubernetes;
— Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards; experience with VictoriaMetrics would be a plus;
— Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization;
— Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management;
— Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs;
— Proficiency in shell command line usage, scripting, and automation tools like Ansible/Terraform;
— Python and bash scripting skills;
Apply
#remote
Post #94
328