MIS audit and stabilization for a clinic network
The vendor walked away and left the code as handed over. In six weeks we closed 152-FZ risks, stopped morning DB outages, and cut the cloud bill nearly in half.
About the project
A federal network of private clinics (dozens of branches, hundreds of doctors, thousands of appointments per day). The web MIS — scheduling, electronic health records, lab and 1C integrations — had been built for about a year and a half by a local contractor. The team dissolved and handed over the sources without support. Peak hours were slow, the database occasionally locked up, and an internal privacy audit raised red flags.
Pain before the engagement
Cloud over budget
Yandex Cloud spend was ≈₽320k/month while the business expected about ₽180k.
Monday morning peak
Opening a patient chart and doctor schedule took more than 8 seconds — clinicians said the system “lags”.
Database outages
About once a month at peak hours locks stalled the DB and front-desk staff could not check patients in.
Risk and money
Double bookings, internal audit findings on 152-FZ / medical secrecy, and broken 1C payment sync that duplicated charges.
Outcomes after 6 weeks
−48%
Yandex Cloud cost: ₽320k → ₽165k / month
0.6 s
p95 chart open at morning peak (was 8.2 s)
152-FZ
follow-up internal audit passed
×4.5
faster CI/CD: build 18 → 4 min
Stack at audit time
Backend
Primary API on FastAPI; admin and users on Django REST Framework. Queues via RabbitMQ, background work on Celery.
Frontend
SPA for reception and doctors on Angular with an Angular Material–based UI kit.
Data & infrastructure
PostgreSQL for operational data, ClickHouse for chief-physician reports. Yandex Cloud: Managed Kubernetes, Managed PostgreSQL, Object Storage. CI on GitLab CE; observability with VictoriaMetrics, Grafana, and EFK.
Audit findings
Critical issues in performance, 152-FZ compliance, and DevOps.
Architecture & performance
Synchronous requests.post() to 1C blocked the FastAPI event loop — a slow 1C stalled the whole API.
N+1 Django ORM on the patient list: 100 rows meant 300+ DB queries.
No PgBouncer: a new connection per request hit max_connections at peak.
ClickHouse “doctor load” report took 40 seconds due to full scans bypassing materialized views.
Security & 152-FZ
Kibana logged full names, insurance policy numbers, and ICD-10 diagnoses in clear text.
RBAC only on the frontend — any doctor could read peers’ schedules and charts via the API.
JWTs in localStorage exposed the app to XSS session theft.
Infrastructure & DevOps
K8s nodes over-provisioned (16 vCPU / 64 GB) with ≤20% utilization.
Homegrown Postgres cron backups without alerts — three weeks without usable copies.
Angular CI builds ~18 minutes with no node_modules cache and legacy Webpack.
What we did
Three two-week stages: security and backups first, then DB and integrations, then right-sizing and CI.
Stage 1 · Security, 152-FZ, backups
FastAPI/Django middleware masking SNILS, passport numbers, and names in logs; Fluentd parsing updated.
Hard RBAC: token doctor_id checked against patient_id (chief physician sees all; doctor — own patients).
JWT moved to HttpOnly / Secure / SameSite=Strict cookies; Angular interceptor sends cookies automatically.
Managed PostgreSQL backups to Object Storage with encryption and 30-day retention; Telegram alert on failure.
Stage 2 · Database & integrations
1C integration via httpx.AsyncClient + RabbitMQ events; Celery worker with exponential backoff — reception UI no longer waits on 1C.
select_related / prefetch_related and targeted RawSQL removed N+1.
PgBouncer in front of Postgres: Too many clients errors gone.
ClickHouse materialized views: doctor report 40 s → 0.8 s.
Stage 3 · Infrastructure & CI/CD
Right-sized nodes to 8 vCPU / 32 GB; backup/log disks moved to network-hdd where appropriate.
esbuild instead of Webpack; Docker layer and node_modules caching in GitLab CI: 18 → 4 minutes.
VictoriaMetrics dashboards and RabbitMQ queue-depth alerts — 1C outages visible before patient complaints.
Key outcomes
Predictable cloud spend
A month of metrics guided VM right-sizing without losing SLA at booking peak.
Stable morning peak
Async 1C, connection pooling, and ORM fixes removed chart lag and lock-driven outages.
152-FZ-ready perimeter
PII masking in logs, server-side RBAC, and secure sessions closed internal audit findings.
1C payment duplicates reduced to zero
Backups created and encrypted on schedule
XSS and horizontal privilege escalation closed