What we monitor
Availability
- HTTP checks against real endpoints, not just a ping
- Status-code and response-body validation
- TLS certificate expiry
- Checks run continuously and are published publicly
Performance
- Response latency per endpoint, tracked over time
- Resource usage per instance, averaged daily
- Degradation alerts before an outage, not after
Security posture
- OS and dependency patching on a defined cadence
- Hardened baseline images
- Least-privilege access, reviewed periodically
- Encrypted transport everywhere, no exceptions
Recovery
- Backup and restore expectations agreed up front
- Restores tested, not assumed
- Documented rollback path for every deploy
Response posture
Something is down
Alerting is automatic and immediate. A human is looking at it, not a ticket queue routing you to a different company.
Something is degraded
We’d rather catch it at “slower than usual” than at “offline”. That’s what the latency tracking is for.
You have a question
One to two business days for anything that isn’t an incident. Reply directly to any email you’ve had from us.