Description
Job Summary:
We are seeking a Tech Lead responsible for managing an Operations Control Center, ensuring service monitoring, observability, and reliability.
Key Highlights:
1. Lead and coordinate control and monitoring activities.
2. Implement advanced monitoring and observability strategies.
3. Promote automation and continuous improvement of service reliability.
**Timestamp SI** is looking for talents who wish to join a dynamic team ambitious to challenge innovation and adopt the best and most recent methodologies and technologies.
**Responsibilities:**
Tech Lead responsible for managing the Operations Control Center (OCC), ensuring implementation of monitoring and observability practices and guaranteeing that all applications, services, and infrastructure are properly monitored.
* Manage and coordinate OCC and monitoring activities.
* Define and implement monitoring and observability strategies for applications, services, and infrastructure.
* Ensure all applications and infrastructure components are integrated into the OCC and covered by appropriate monitoring mechanisms.
* Define metrics, indicators, dashboards, and operational alerts, ensuring their relevance and effectiveness.
* Promote automation and scripting of operational tasks and monitoring processes.
* Monitor events, incidents, and problems, ensuring timely analysis, escalation, and resolution.
* Ensure compliance with agreed service levels and incident and event management processes.
* Support definition and evolution of operational processes aligned with ITIL best practices.
* Collaborate with development, infrastructure, and operations teams to identify risks and continuously improve service reliability.
* Produce technical and operational documentation, reporting performance, availability, and monitoring quality indicators.
**Required Profile:**
* Bachelor’s degree in Computer Engineering, Software Engineering, or related field.
* Ability to collaborate with technical and business teams, translating operational requirements into monitoring solutions.
* Strong written and verbal communication skills.
* Experience in technical leadership and/or management of control, operations, and monitoring centers.
* Experience implementing monitoring and observability solutions for technology services.
* Knowledge of IT operations automation and scripting.
* Experience configuring dashboards, metrics, indicators, and operational alerts.
* Knowledge of incident, event, problem, and service level management.
* Familiarity with ITIL methodologies and best practices.
* Experience analyzing application and infrastructure availability, performance, and capacity.
**Preferred Qualifications:**
* Experience with monitoring, observability, event management, and alerting tools.
* Knowledge of cloud environments, infrastructure, networks, operating systems, and databases.
* Experience with automation and scripting tools, particularly Python, PowerShell, or Bash.
* Understanding of logging, tracing, metrics, APM, and API monitoring concepts.
* Experience in critical production environments with high availability requirements.
* Experience defining and tracking SLAs, OLAs, and operational KPIs.