Building an LMS looks simple at the beginning.
You need users, classes, payments, attendance, recordings, homework, notifications, and reports.
Then you move into production.
Suddenly, you are dealing with authentication, permissions, concurrent requests, payment verification, background jobs, database indexes, cloud costs, logging, deployments, backups, and uptime.
I have been working on a production LMS used across web and mobile, where students attend live classes, tutors manage academic activities, payments are processed online, and several services need to work together reliably.
This post covers the technical decisions that worked, the problems that became more important as the system grew, and what I would change if I started the platform again today.
1. The Architecture
Our main stack includes:
- Node.js
- NestJS
- MongoDB
- AWS
- Cloudflare
- Docker
- React / Next.js
- Flutter
At a high level, the architecture looks like this:
Students / Tutors
|
+---------+---------+
| |
v v
Web App Mobile App
React / Next.js Flutter
| |
+---------+---------+
|
v
Cloudflare
|
v
Node.js / NestJS API
|
+----------------+----------------+
| | |
v v v
MongoDB AWS Services External APIs
|
+--------+--------+
| | |
v v v
S3 SES CloudFront
This architecture has worked well for us.
But the interesting part is not the technology itself.
The important part is how you design the system around it.
2. Node.js Was Not the Problem
One question I often hear is:
Can Node.js handle a serious production system?
For this type of workload, yes.
A typical LMS performs a large amount of I/O work.
Examples include:
- Reading student profiles
- Fetching timetables
- Loading class information
- Processing attendance
- Checking enrollments
- Verifying payments
- Loading recordings
- Sending notifications
- Calling external APIs
- Reading and writing database records
Node.js handles this type of workload well.
The bigger performance problems usually come from:
- Poor database queries
- Missing indexes
- Too many database round trips
- Unnecessary API calls
- Blocking work inside request handlers
- Poor background job design
- Large payloads
- Weak caching strategies
A slow API does not automatically mean your runtime is slow.
Consider a query like this:
await Student.find({
instituteId,
status: "ACTIVE",
classIds: classId,
});
It looks harmless.
But as the collection grows, the wrong indexing strategy can make this request increasingly expensive.
A suitable index might look like:
studentSchema.index({
instituteId: 1,
status: 1,
classIds: 1,
});
Of course, indexes should be designed around your actual query patterns and verified using query execution statistics.
The lesson is simple:
Fix the real bottleneck before replacing the entire technology stack.
Moving from Node.js to another runtime will not fix an unindexed database query.
3. MongoDB Worked Well, but Schema Design Still Matters
MongoDB is easy to start with.
That flexibility can also create problems if you treat schema design as optional.
It is not optional.
Imagine storing every attendance record inside a student document:
{
"studentId": "ST001",
"attendance": [
{},
{},
{},
{}
]
}
This might look convenient when the system is small.
It becomes harder to manage when students accumulate large amounts of historical data.
For frequently growing operational data, I prefer dedicated collections.
For example:
students
tutors
classes
enrollments
attendance
payments
recordings
homework
notifications
audit_logs
This makes it easier to:
- Index the data correctly
- Query records independently
- Paginate history
- Archive old records
- Maintain auditability
- Avoid endlessly growing documents
MongoDB gives you flexibility.
It does not remove the need for database architecture.
4. Attendance Is More Complex Than It Looks
An LMS attendance flow sounds simple:
Student joins class
|
v
Mark present
Production requirements make it much more complicated.
You may need to know:
- When the class started
- When the student joined
- Whether the student joined late
- How many minutes late they were
- Whether the tutor joined
- Whether the student disconnected and rejoined
- Whether an attendance record already exists
- Whether the class session is actually valid
- Whether the student is enrolled in the class
Now concurrency matters.
Two requests arriving almost at the same time should not create two attendance records.
This pattern is risky:
const attendance = await Attendance.findOne({
studentId,
classId,
sessionId,
});
if (!attendance) {
await Attendance.create(data);
}
There is a gap between the read and the write.
Two concurrent requests can both see no record and both attempt to create one.
I prefer protecting the invariant at database level:
attendanceSchema.index(
{
studentId: 1,
classId: 1,
sessionId: 1,
},
{
unique: true,
},
);
Then the application handles duplicate key conflicts safely.
Application checks are useful.
Database constraints are stronger.
5. Payments Must Be Idempotent
Never assume a payment callback will arrive only once.
Payment providers can retry callbacks.
Network failures happen.
Your API can retry.
Users can refresh pages.
Workers can retry failed jobs.
Imagine this flow:
Payment successful
|
v
Store payment
|
v
Activate enrollment
|
v
Create invoice
|
v
Update balance
If the same transaction is processed twice, the consequences can be serious.
Every payment should have a unique reference from the payment provider or your own transaction system.
A basic duplicate check may look like:
const existingPayment = await Payment.findOne({
gatewayTransactionId,
});
if (existingPayment) {
return existingPayment;
}
Then enforce uniqueness at database level:
paymentSchema.index(
{
gatewayTransactionId: 1,
},
{
unique: true,
},
);
For more complex payment flows, I also want the transaction state machine to be explicit.
For example:
PENDING
|
+----> PAID
|
+----> FAILED
|
+----> CANCELLED
A confirmed payment should not accidentally move backwards because a delayed callback arrived later.
Idempotency is one of the most important patterns in any system that handles money.
6. Do Not Put Everything Inside the API Request
A common early architecture looks like this:
await createEnrollment();
await sendEmail();
await sendSMS();
await sendNotification();
await generateInvoice();
await updateAnalytics();
return response;
It works during development.
But now the user is waiting for every downstream service.
If the SMS provider takes four seconds to respond, your endpoint may also take four extra seconds.
If the email provider fails, should the entire enrollment fail?
Usually, no.
A better architecture is:
API Request
|
v
Validate Request
|
v
Critical Database Operation
|
v
Return Response
|
v
Background Queue
|
+--> Email
|
+--> SMS
|
+--> Push Notification
|
+--> Analytics
|
+--> Non-critical processing
The synchronous request should handle work required to guarantee the main business operation.
Non-critical work should move to background processing.
This improves:
- API latency
- Reliability
- Retry handling
- Failure isolation
- User experience
It also gives you a better place to manage external provider failures.
7. Cloudflare Became an Important Layer
Cloudflare is not only DNS.
For a public platform, it can provide several useful controls before traffic reaches your application.
Examples include:
- DNS
- TLS
- CDN
- DDoS protection
- Web Application Firewall
- Rate limiting
- Bot protection
- Caching
The request path becomes:
Internet
|
v
Cloudflare
|
v
AWS / Application Infrastructure
|
v
Node.js API
Rate limiting is especially important for sensitive endpoints.
Examples:
POST /auth/login
POST /auth/password-reset
POST /auth/send-otp
POST /register
POST /payments/verify
These endpoints should not accept unlimited requests from the same source.
The exact rate limits depend on the endpoint and your users.
The key point is to protect expensive and sensitive operations before abuse becomes a production issue.
8. AWS Cost Is Also an Architecture Problem
AWS gives you many ways to solve the same problem.
That flexibility is useful.
It can also become expensive when services are added without understanding how billing behaves at scale.
I mentally separate infrastructure into three categories.
Critical
Services required for the product to function.
Compute
Database
Object storage
Backups
Networking
Operational
Services required to operate the product reliably.
Monitoring
Logging
Alerts
CI/CD
Security tooling
Optional
Services that provide convenience but are not always required immediately.
Before adding a new managed service, I ask:
- What exact problem does this solve?
- Can our existing infrastructure solve it?
- What happens when usage increases 10x?
- What is the expected monthly cost?
- What does data transfer cost?
- How difficult is it to migrate away later?
- Is the operational saving worth the additional cost?
Cloud architecture is also cost architecture.
A solution is not scalable if the business cannot afford the scaling curve.
9. Logging Changed How We Debug Production
During development, this is common:
console.log(error);
That is not enough for a production platform.
Logs need context.
This log is not very useful:
Payment failed
A structured event is much easier to investigate:
{
"level": "error",
"event": "PAYMENT_VERIFICATION_FAILED",
"studentId": "ST001",
"transactionId": "TX12345",
"provider": "payment_gateway",
"requestId": "REQ-8F23A1",
"timestamp": "2026-08-10T04:00:00Z"
}
Now the operations team can answer useful questions.
- What happened?
- When did it happen?
- Which user was affected?
- Which transaction failed?
- Which request triggered it?
- Which external service was involved?
I also like carrying a request or correlation ID across services.
For example:
Client Request
|
| requestId: REQ-8F23A1
v
API
|
+--> Database log
|
+--> Payment log
|
+--> Background job
|
+--> Notification log
When something breaks, one ID can help trace the entire flow.
10. Authentication and Authorization Are Different Problems
Authentication answers:
Who is this user?
Authorization answers:
What is this user allowed to do?
They should not be treated as the same thing.
An LMS can have roles such as:
Student
Tutor
Academic Coordinator
Student Consultant
Finance Staff
Administrator
Super Administrator
All of these users may be authenticated.
They should not have the same capabilities.
A system built entirely around checks like this can become difficult to maintain:
if (user.role === "ADMIN") {
// allow action
}
As the product grows, I prefer permissions that represent capabilities.
For example:
class.create
class.view
class.update
attendance.view
attendance.update
payment.view
payment.verify
recording.view
recording.manage
student.view
student.update
user.permission.manage
Roles then become collections of permissions.
Tutor
class.view
attendance.view
homework.create
homework.grade
recording.view
Finance Staff
payment.view
payment.verify
invoice.view
This model scales better when responsibilities change.
It also reduces the temptation to give a user a powerful role just because they need one extra action.
11. Validate Data at the API Boundary
Frontend validation improves user experience.
It is not a security boundary.
Anyone can call your API directly.
That means the backend must validate incoming data independently.
For example:
const schema = Joi.object({
classId: Joi.string().required(),
studentId: Joi.string().required(),
amount: Joi.number().positive().required(),
});
I prefer this flow:
Request
|
v
Validation
|
v
Authentication
|
v
Authorization
|
v
Controller
|
v
Service
|
v
Database / External Services
Your service layer should receive structurally valid data.
This keeps business logic cleaner and reduces defensive checks throughout the codebase.
12. Keep Controllers Small
One backend pattern I strongly prefer is:
Controller
|
v
Service
|
v
Repository / Database
Controllers should coordinate HTTP concerns.
They should not become the entire application.
This becomes difficult to maintain:
@Post()
async create() {
// validate request
// query student
// check enrollment
// verify payment
// create enrollment
// send notification
// send email
// create audit log
// update analytics
}
A cleaner controller might look like:
@Post()
async create(@Body() dto: CreateEnrollmentDto) {
return this.enrollmentService.create(dto);
}
Then the service owns the business process.
For example:
@Injectable()
export class EnrollmentService {
async create(dto: CreateEnrollmentDto) {
const student = await this.getStudent(dto.studentId);
await this.assertCanEnroll(student, dto.classId);
const enrollment = await this.createEnrollment(dto);
await this.queuePostEnrollmentTasks(enrollment);
return enrollment;
}
}
This improves:
- Testability
- Readability
- Reuse
- Debugging
- Maintenance
The service should still be kept focused.
A single 2,000-line service is not better than a 2,000-line controller.
13. Observability Is More Than CPU and RAM
A server showing:
CPU: 20%
RAM: 40%
does not mean the application is healthy.
Those metrics are useful, but they are only part of the picture.
For a production LMS, I care about metrics such as:
- API response time
- API error rate
- Database latency
- Slow queries
- Failed background jobs
- Queue depth
- Payment verification failures
- Login failures
- SMS failures
- Email failures
- External API latency
- Application restarts
- Memory growth
- Disk usage
- Availability
I also want alerts based on useful thresholds.
For example:
5xx rate > normal threshold
Payment callback failures increasing
Queue backlog continuously growing
Database latency suddenly increasing
Application restarting repeatedly
Storage nearing capacity
The goal is simple.
I want the engineering team to know about an important problem before students or tutors need to report it.
14. Background Jobs Need Retry Rules
Moving work to a queue does not automatically make it reliable.
You also need to decide what happens when a job fails.
For example:
Send SMS
|
v
Provider timeout
|
v
Retry
|
v
Provider timeout
|
v
Retry later
Not every job should retry forever.
A useful policy might look like:
Attempt 1: immediately
Attempt 2: after 30 seconds
Attempt 3: after 2 minutes
Attempt 4: after 10 minutes
Final: mark failed and alert if important
This is a simplified example.
The correct policy depends on the operation.
Payment processing, email delivery, recording synchronization, and analytics may all need different retry behaviour.
You should also make jobs idempotent whenever possible.
A retried job should not accidentally create duplicate data.
15. Audit Logs Are Worth Adding Early
Production systems eventually reach a point where somebody asks:
Who changed this?
You need an answer.
For important administrative actions, I want an audit event containing data such as:
{
"actorId": "USR001",
"action": "STUDENT_STATUS_UPDATED",
"entityType": "student",
"entityId": "ST001",
"previousValue": {
"status": "ACTIVE"
},
"newValue": {
"status": "SUSPENDED"
},
"timestamp": "2026-08-10T04:30:00Z"
}
Examples of operations worth auditing include:
- Payment changes
- Refund actions
- Student status changes
- Enrollment changes
- Role changes
- Permission changes
- Tutor changes
- Manual attendance changes
- Important configuration changes
Application logs tell you what the system did.
Audit logs tell you what users and administrators changed.
They solve different problems.
16. Security Needs Multiple Layers
There is no single "security feature" that makes a platform secure.
For a production LMS, I think about security across multiple layers.
Edge
Cloudflare
WAF
Rate limiting
DDoS protection
Bot controls
Application
Authentication
Authorization
Input validation
Secure session/token handling
Idempotency
Safe error handling
Data
Least privilege
Encryption
Backups
Database access controls
Audit logs
Infrastructure
Network restrictions
Secrets management
Patch management
Container security
Monitoring
Development Process
Dependency scanning
Code review
CI/CD controls
Secret scanning
Environment separation
Security becomes much easier when these controls are part of the architecture from the beginning.
17. Separate Development, Staging, and Production
One lesson that becomes obvious as a system grows is that environments need clear boundaries.
At minimum, I want:
Local
Development / Staging
Production
Production credentials should never be casually used during development.
The same applies to:
- Databases
- Payment gateway credentials
- SMS credentials
- Email credentials
- Storage buckets
- API keys
- Webhooks
Environment separation protects real users and real data.
It also makes deployment testing much safer.
18. Backups Are Only Useful if You Can Restore Them
"Backups enabled" is not enough.
The real question is:
Can we restore the system successfully when we need to?
A production backup strategy should consider:
- Database backups
- File/object backups
- Retention periods
- Restore procedures
- Restore testing
- Recovery time objectives
- Recovery point objectives
If your team has never tested a restore, you do not fully know whether your backup process works.
Recovery should be treated as an engineering workflow, not just a checkbox.
19. What I Would Change If I Started Again
If I rebuilt the LMS today, I would make several decisions earlier.
1. Design background jobs from day one
Email, SMS, notifications, recordings, analytics, and other non-critical processing should not block normal API requests.
2. Add structured logging immediately
Debugging production without useful logs wastes engineering time.
3. Add correlation IDs
Tracing one request across the API, workers, database operations, and external services becomes much easier.
4. Design permissions before adding many roles
Role-only authorization becomes harder to maintain as the organization grows.
5. Create indexes from real query patterns
Do not add indexes randomly.
Measure how the application actually reads data.
6. Make payment operations idempotent
Anything involving money must handle duplicate callbacks and retries safely.
7. Add audit logs for important changes
You need to know who changed critical data.
8. Treat infrastructure cost as an engineering metric
A technically impressive architecture that costs more than the business can support is not a good architecture.
9. Test failure paths
Do not test only the happy path.
Test what happens when:
MongoDB is slow
SMS provider fails
Email provider fails
Payment callback is duplicated
Worker crashes
External API times out
User sends the same request twice
10. Keep the architecture understandable
A small team should not need 20 services just because microservices are popular.
Complexity must solve a real problem.
20. Would I Still Choose Node.js?
For this type of system, yes.
I would still be comfortable choosing a stack such as:
Node.js
NestJS
MongoDB or PostgreSQL
AWS
Cloudflare
Docker
Node.js provides good development speed and fits I/O-heavy workloads well.
Would Spring Boot also work?
Yes.
Would Django work?
Yes.
The framework is rarely the main reason a production platform succeeds or fails.
The surrounding engineering decisions matter more.
A badly designed Spring Boot application can fail.
A badly designed Node.js application can fail.
A well-designed version of either can handle serious production workloads.
21. MongoDB or PostgreSQL If I Started Today?
This is one area where I would spend more time on the data model before choosing.
MongoDB works very well when the application's data naturally benefits from document-oriented modelling and flexible structures.
PostgreSQL becomes attractive when the system has increasingly relational workflows such as:
- Students
- Classes
- Enrollments
- Invoices
- Payments
- Attendance
- Academic periods
- Permissions
- Reporting
For a new LMS, I would evaluate these relationships carefully before making the database decision.
The correct answer is not:
MongoDB is always better
or:
PostgreSQL is always better
The better question is:
Which data model makes the core business rules easier to enforce, query, report on, and maintain?
That is the database decision I care about now.
22. The Architecture I Prefer Today
If I were starting a similar platform today, I would keep the first production architecture relatively simple.
Cloudflare
|
v
Web / Mobile Clients
|
v
API Load Balancer
|
v
Node.js / NestJS API
|
+--------------+--------------+
| | |
v v v
Database Queue/Cache AWS S3
| |
| v
| Background Workers
| |
| +-------+-------+
| | | |
| v v v
| SMS Email Other APIs
|
v
Backups / Monitoring
I would not introduce additional architectural complexity until the product demonstrates a clear need for it.
A modular monolith can be a very good architecture.
You can still have:
- Clear module boundaries
- Background workers
- Independent queues
- Strong observability
- Horizontal scaling
- Good security controls
without immediately splitting everything into microservices.
Final Thoughts
Building production software changed how I evaluate technology.
I care less about questions like:
Which framework is fastest?
I care more about:
Can we debug it?
Can we secure it?
Can we scale it?
Can we recover from failure?
Can another developer maintain it?
Can we observe it?
Can the business afford to operate it?
Those questions matter much more once real students, tutors, classes, and payments depend on the platform.
The biggest lesson for me has been this:
Production architecture is not about choosing the most powerful technology. It is about making good engineering decisions around the technology you already have.
If you were building an LMS today, which backend would you choose?
Node.js + MongoDB, Node.js + PostgreSQL, Spring Boot + PostgreSQL, Django, or something else?
I would especially like to hear from developers who have operated these stacks in production.
Top comments (0)