Senior Backend Engineer · Fintech SaaS · Platform Architecture

FinFlow SaaS Platform
B2B Fintech · SaaS

The client was running a 7-year-old monolithic accounting system that couldn't hold up against their growing enterprise customer base. I designed and built a replacement from scratch: a multi-tenant B2B SaaS platform with six independently deployable NestJS microservices, event-driven reconciliation via Redis Streams, full RBAC with tenant data isolation, observability from day one, and a GitHub Actions CI/CD pipeline with zero-downtime blue/green deployments. The result was a 12x increase in transaction capacity, 99.9% uptime SLA, and deploy cycles cut from 3 days to 12 minutes.

The Challenge

A 7-year-old monolith at its breaking point

The client operated a monolithic accounting system that had accumulated 7 years of incremental growth without architectural oversight. What started as a small internal tool had become the financial backbone for dozens of enterprise clients — and it was failing under the weight. During peak transaction windows, the system buckled. Reconciliation was an entirely manual operation that consumed 8+ hours per week. There was no observability infrastructure, no access control beyond basic login, and every deployment involved a manual multi-day coordination process that blocked the entire team.

  • Monolith failing under 500+ concurrent users — 14.2 hours of downtime in a single month
  • Manual reconciliation consuming 8+ hours per week with no automation path
  • No role-based access control — flat permission model created serious compliance exposure
  • Zero observability: no structured logging, no alerting, no distributed tracing
  • 3-day deployment cycle with manual handoff steps and no rollback capability

legacy_system.statusCritical
78%
Error rate
4.2s
Avg latency
↓12%
Throughput
error spikes · last 8h
Downtime this month14.2h
Failed reconciliations342
Support tickets open89
The Solution

Six microservices. Event-driven. Deployed in 12 minutes.

1
Domain-driven microservice decomposition - Six bounded contexts extracted from monolith:
  • Services: Transactions, Reconciliation, Reporting, Auth, Notifications, Admin
  • Each service owns its own database (no cross-service schema dependencies)
  • Foundational decision enabling independent scaling & faster releases
2
Event-driven reconciliation via Redis Streams - Automated what was an 8-hour manual weekly process:
  • Redis Streams with consumer groups (guaranteed delivery + retry)
  • Idempotency keys prevent double-processing
  • Distributed locks prevent race conditions
  • 8-hour weekly manual process → fully automated
3
Multi-tenant RBAC with data isolation - Secure, compliant, SOC2-ready access control:
  • Organisation-scoped role-based access control
  • Permissions enforced at API Gateway + service layer
  • Tenant data isolation at database query level (no escapes)
  • SOC2 audit logging with correlation IDs on every write
4
Observability architecture from day one - Full visibility into system behavior:
  • Structured JSON logging across all six services with correlation IDs
  • CloudWatch dashboards: four golden signals per service
  • PagerDuty integration for P1 alerts
  • Runbooks written for every failure scenario before production
5
CI/CD with zero-downtime blue/green deployments - Automated, safe, reliable releases:
  • GitHub Actions → Docker → AWS ECS blue/green
  • Automated health check validation gates deployments
  • Rollback triggers automatically on health check failure
  • Deploy time: 3-day manual process → 12 minutes with zero human intervention
FinFlow SaaS Architecture
Enterprise ClientsWeb Dashboard · API ConsumersMulti-Tenant IsolationJWT + Tenant context · RBACJWT + tenant headerAPI GatewayRate limiting · RoutingAuth ServiceRBAC · Role enforcementvalidated + routedTransactions1M+ txns/dayIdempotency keysReconciliationRedis Streams driven8hr manual → autoReportingAsync aggregationRead-only DB replicaNotificationsEvent-driven triggersEmail · SMS · PushAdminTenant managementAudit logsDistributed LockRedis · race preventionIdempotency enforcedevents · consumer groupsREDIS STREAMS · EVENT BUSConsumer groups · Guaranteed delivery · Retry on failureIdempotency keys · Distributed locks · Dead-letter handlingTxn + Recon DBPostgreSQL · ownedReporting DBRead replica · no cross-svcAuth + Admin DBNo shared schemasGitHub Actions CI/CD · Blue/Green ECS deployments · Zero-downtimeDeploy time: 3 days → 12 minutes · 99.9% uptime SLA
Redis Streams: Event-Driven Reconciliation
STREAM PRODUCERSTransaction ServiceXADD stream:transactions *Ledger ServiceXADD stream:ledger *REDIS STREAMS · CONSUMER GROUPstream:transactionsRedis · append-only · ACKstream:ledgerRedis · append-only · ACKConsumer GroupXREADGROUP NOACK · 3 workersReconciliation Workermatch · validate · settleRECONCILIATION OUTCOMESMatchXACK stream · mark settledPostgreSQL ledger updateMismatchDLQ → alert + retryops on-call notificationAudit Logimmutable · PostgreSQLRedis Streams · consumer groups · XACK · race-condition-safe · audit trail
Results & Impact

Numbers that moved the business

After a 14-week build and phased migration, the new platform went live. Here's what changed.

1M+
Transactions/day
Up from 80K - 12× capacity increase
99.9%
Uptime SLA
Down from 14.2h downtime/month to ~43 min
40%
Faster load time
From 4.2s avg to 250ms at P95
0h
Manual reconciliation
8 hours/week fully automated
3×
Engineering velocity
Deploy time from 3 days to 12 minutes
SOC2
Compliance-ready
Audit logging + RBAC from day one
My Role

End-to-end ownership

⚛
Backend Architecture & Development
Designed and built the entire frontend (Next.js) and all backend microservices (NestJS), including auth, transactions, and reporting APIs.
⬡
System Architecture
Designed the microservices decomposition strategy, event-driven messaging layer, database schema, and multi-tenant data isolation approach.
▲
DevOps & Deployment
Set up the entire AWS infrastructure using ECS, configured GitHub Actions CI/CD, and implemented CloudWatch alerting and log aggregation.
Live Product

What the dashboard looks like today

finflow.dashboard - productionAll Systems Nominal
99.9%
Uptime SLA
247ms
Avg Latency P95
1.2M
Txns today
transaction throughput · stable 📈
Reconciliation runs today12 / 12 completed ✓
Failed transactions0
Active tenants47
Engineering Decisions

Architectural choices, tradeoffs, and what nearly broke production

Monolith to microservices, but how many?

Why six services and not three or ten?

Six emerged from the business domain boundaries, not an arbitrary number. Transactions, Reconciliation, Reporting, Auth, Notifications, and Admin are each independently scalable and deployable. Beyond that, coordination overhead exceeds benefit. Too few and you're back to interdependencies that defeat the point.

Why Redis Streams instead of traditional message queue?

Wouldn't RabbitMQ or Kafka be more robust for a financial system?

In fintech, you want both reliability and simplicity. Redis Streams gave us guaranteed message delivery with consumer groups, automatic acknowledgment, and dead-letter handling. For a 1M transaction/day system, it's more than sufficient. Kafka added operational complexity we didn't need. Redis we already had for caching.

How do you prevent reconciliation race conditions?

What happens if two reconciliation jobs run simultaneously?

Idempotency keys on every transaction. Even if a reconciliation job re-processes the same batch, the second run produces the same result without double-counting.

The reconciliation service locks the batch it's processing. Subsequent jobs skip locked batches and process only unlocked work. If a job crashes mid-process, a heartbeat mechanism releases the lock after 10 minutes.

Why structured logging from day one?

Wouldn't basic console logs be enough to get started?

No. In a fintech system, you need to trace every transaction from entry to exit. If a transaction disappears or records incorrectly, you need forensics: which service touched it, what decision was made, what external API returned, and when. We structured logging on day one so that when incidents happen (and they will), you have answers in minutes, not hours.

THE TRADEOFFS

Cost · Risk · Tradeoffs

Estimated values based on production metrics and monitoring data

Cost Impact
small to medium
Team size
Microservices require more infrastructure overhead and coordination. Traded monolith simplicity for independent deployment autonomy and scaling flexibility.
medium increase
Troubleshooting complexity
Debugging across six services requires better observability. Structured logging and distributed tracing overhead is worth it for production confidence.
Risk Reduction
improved 100%
Deployment safety
No more 3-day deployments with manual handoffs. Blue-green deployments and automated rollback mean safer, more frequent releases with zero downtime.
Technical Tradeoff
99.99% guaranteed
Reconciliation accuracy
Idempotency and distributed locks prevent race conditions but add latency. Trade accepted: reconciliation runs 2-3 seconds slower, but produces correct results every time.
Before vs After

Transformative Results: Before vs After

✕Before
  • Legacy monolith under loadPerformance bottleneck
  • Manual reconciliation process8+ hrs/week
  • No role-based access controlSecurity & compliance risk
  • Zero observability stackNo monitoring
→
✓After
  • 6-service microarchitecture1M+ txns/day
  • Automated reconciliationNear-zero manual work
  • Full RBAC implementationSecure & compliant
  • Full observability stackReal-time monitoring
⇪Business Impact
1M+ transactions/day
8+ hrs saved weekly
0 downtime deploys
Tech Stack

What was used

FrontendNext.js · React
BackendNestJS · Node.js
DatabasePostgreSQL · Redis
CloudAWS ECS · S3 · CloudWatch
DevOpsDocker · GitHub Actions
LanguageTypeScript

Building a fintech platform or scaling a legacy system?

I design backend systems that handle real financial workloads - multi-tenancy, reconciliation, compliance, and the CI/CD discipline to ship them safely. Let's talk about your architecture.