Site Reliability & Monitoring
Selected work demonstrating this service.
- Developed and launched the company's first observability dashboard, providing real‑time system performance insights and data visualization on the large office TV.
- Designed, deployed, and maintained 10 PostgreSQL and MS SQL servers on Ubuntu Linux VPS, ensuring optimal server performance and reliability.
- Mitigated operational risks by implementing a monitoring dashboard using Grafana and Prometheus, improving system reliability.
- Administered 40 websites on Ubuntu Linux hosting servers with Apache and Nginx, ensuring high availability and performance.
- Architected, developed, implemented, supported infrastructure, data processing, and the map application for 2 years non‑stop without any weekends, holidays, or vacations, 10‑14 hours a day.
- Implemented startup mail‑provider validation across three providers (SendGrid, SMTP2GO, Azure ACS) with environment‑aware behaviour — logging in development, failing the boot in staging and production.
- Owned end‑to‑end deployments of the platform to Azure, managing releases across development, staging and production environments.
- Provided round‑the‑clock 24/7 infrastructure support for an IPTV/OTT streaming platform, administering ~1,000 servers plus client‑owned systems for global customers in China, the US and Germany.
- Ensured uninterrupted delivery of IPTV streaming signals between suppliers and clients, monitoring and maintaining the streaming network and IP telephony around the clock.
- Planned and implemented new infrastructure functionality for internal and external systems, building solutions durable enough to still run years later with minimal change.