Traditional Chinese version: [[content/posts/operations-hospital-collaboration-guide]]
1. Purpose
This guide is intended for the TAMS Backend, Hospital Integration, RMF/Robot Integration, QA, and Operations teams. It describes the actual boundaries of the current Modular Monolith, how Operations and Hospital communicate, and the rules that all contributors must follow when changing the codebase.
The system is still deployed as one NestJS process, one npm package, and one versioned release. Modules communicate through NestJS dependency injection and in-process method calls. There are no internal HTTP calls, message brokers, or additional worker frameworks between modules.
The primary design principle is:
Hospital decides what the hospital workflow requires. Operations guarantees how robot tasks are executed safely. Adapters handle how the system communicates with external technologies.
2. Quick Placement Guide
| Requirement | Owner/location | Examples |
|---|---|---|
| General task, robot, or facility rules | src/operations/<feature> | Task cancellation, dispatch, robot availability, charging-station lookup |
| Background workflows spanning multiple Operations features | src/operations/workflows | Queue Worker, Robot Keeper |
| CSH/CMP/HIS or hospital-specific workflows | src/csh-hospital | Emergency return to pharmacy, employee sync, medication lookup |
| HTTP validation and response mapping | src/product-api, src/csh-hospital/api, or the existing controller layer | DTOs, guards, Swagger decorators |
| MongoDB, Redis, RMF, Axios, ROS, or Socket.IO | src/adapters or the existing infrastructure layer | Repositories, RMF adapters, robot-control adapters |
| Authentication, JWT, RBAC, sessions, and audit | src/access | Actor, permissions, user management |
| Pure types shared by multiple top-level modules | src/domain | Types that do not belong to one Operations feature |
If a requirement contains a hospital-name condition such as if (hospital === 'CSH'), do not add it directly to Operations. First determine whether configuration can express the difference. Introduce a small, explicit cross-module contract only when the behavior is genuinely different.
3. System Architecture
flowchart LR
ProductAPI[Product API controllers]
HospitalAPI[Hospital API controllers]
HospitalWF[CSH Hospital workflows]
subgraph Operations[Operations Module]
Task[Task feature]
Robot[Robot feature]
Fleet[Fleet feature]
Facility[Facility feature]
Settings[Settings feature]
CrossWF[Queue Worker / Robot Keeper]
end
SharedPorts[Operations shared ports]
LocalPorts[Feature-local outbound ports]
Adapters[Mongo / Redis / RMF / Robot control / WebSocket adapters]
ProductAPI -->|feature public entry| Task
ProductAPI -->|feature public entry| Robot
ProductAPI -->|feature public entry| Fleet
ProductAPI -->|feature public entry| Facility
ProductAPI -->|feature public entry| Settings
HospitalAPI --> HospitalWF
ProductAPI -->|medicine-cart endpoint| HospitalWF
HospitalWF -->|TaskProgressionService public entry| Task
HospitalWF -->|EmergencyOperationsPort| SharedPorts
SharedPorts -->|implemented by emergency facade| CrossWF
HospitalWF -->|SiteOperationsConfig implementation| SharedPorts
Task --> LocalPorts
Robot --> LocalPorts
Fleet --> LocalPorts
Facility --> LocalPorts
Settings --> LocalPorts
CrossWF --> Task
CrossWF --> Robot
LocalPorts -->|implemented by| Adapters
SharedPorts -->|events/config adapters| Adapters3.1 Dependency Direction
The normal dependency direction is:
API / Hospital workflow
↓
Operations application
↓
Domain + Ports
↑
Adapters implement PortsOperations must not import a concrete CSH Hospital implementation. Hospital obtains general task capabilities through the task feature’s public entry and executes complete cross-feature Operations use cases through EmergencyOperationsPort. Operations returns neutral outcomes and never calls back into Hospital.
4. Feature-First Operations Structure
src/operations/
├── task/
│ ├── domain/
│ ├── application/
│ │ └── ports/
│ ├── public/
│ └── task.module.ts
├── robot/
│ ├── domain/
│ ├── application/
│ │ └── ports/
│ ├── public/
│ └── robot.module.ts
├── fleet/
├── facility/
├── settings/
├── workflows/
├── ports/
└── operations.module.ts4.1 Responsibilities Within a Feature
domain
- Contains business entities, enums, records, value objects, and pure rules.
- Must not depend on NestJS, Swagger, MongoDB/Mongoose, Redis, Axios, ROSLIB, or Socket.IO.
- Must not contain HTTP requests, Axios responses, Mongoose documents, or raw RMF payloads.
- IDs crossing domain, application, and port boundaries must be strings.
application
- Executes a concrete use case and coordinates domain logic with ports.
- May use NestJS dependency injection, but must not import concrete adapters, API DTOs, or Hospital implementations.
- Must not construct Axios, RMF, or Redis wire payloads directly.
application/ports
- Defines the external capabilities required by the owning feature, such as repositories, runtime-state stores, RMF commands, or robot control.
- Ports should be designed around use-case needs rather than generic CRUD repositories.
- Port commands and results must be transport-neutral.
public
- Provides the public import entry used by APIs, Hospital, CLIs, and inbound adapters.
- The current
public/index.tsis a controlled export barrel, not an additional proxy or separate service instance. - Before exporting a new service, verify that an external module genuinely needs it. Do not re-export every application service by default.
*.module.ts
- Registers and exports providers owned by the feature.
- Feature modules do not contain HTTP controllers.
OperationsModuleonly composes feature modules and cross-feature workflows.
4.2 Feature Ownership
| Feature | Responsibilities | Main public capabilities |
|---|---|---|
| Task | Task history, task libraries, queue intent, dispatch, cancellation, continuation | TaskApplicationService, TaskRequestService, TaskCancelService |
| Robot | Runtime state, availability, navigation, snapshots, video | RobotApplicationService, RobotVideoService, UpdateRobotStateService |
| Fleet | Fleet CRUD, fleet commands, Fleet Manager discovery | FleetApplicationService, FleetService |
| Facility | Maps, doors/lifts, devices, charging lookup | DeviceApplicationService, MapFloorService |
| Settings | System settings | SystemSettingsApplicationService |
TaskQueueWorker and RobotKeeperWorker depend on several features. They therefore live in operations/workflows and are registered by OperationsModule.
5. CSH Hospital Module
src/csh-hospital/
├── api/ Hospital HTTP controllers
├── config/ CSH/site configuration and typed access
├── dto/ Hospital transport DTOs
├── integrations/ CSH/CMP/HIS HTTP integrations
├── ports/ Hospital-owned persistence/integration ports
├── workflows/ Hospital business workflows
└── csh-hospital.module.tsThe Hospital module is responsible for:
- Interpreting CMP/HIS/CSH requests and event semantics.
- Employee, medication, and hospital-data integrations.
- Deciding which hospital workflow an emergency requires.
- Using Operations capabilities to cancel, create, pause, or continue tasks.
- Providing site and hospital configuration without forcing Operations to inspect hospital names.
The Hospital module must not:
- Build RMF payloads or call Open-RMF directly.
- Operate on Redis keys or MongoDB collections directly.
- Reimplement queue leasing, robot availability, or dispatch idempotency.
- Bypass Operations cancellation and dispatch safety rules.
6. Communication Between Operations and Hospital
6.1 Hospital Calls Operations
A Hospital controller first calls a Hospital workflow. The workflow then uses task and robot capabilities provided by Operations.
For an emergency event:
POST /cmp/messages/hospitalreceivesReturnToErPharmacy.HospitalMissionControlService.requestForActiveTasks()first persists fleet-widetaskAcceptancePausedAtand a newtaskAcceptancePauseVersionin MongoDB.- It obtains active task contexts through
EmergencyOperationsPort. If MongoDB or Redis cannot be read, the flow fails closed and the persisted fleet pause remains active. - It broadcasts
EMERGENCY_STARTEDto every robot, not only active robots. - For every active task, it decides whether:
- The robot is moving: cancel the original task and create a return-to-pharmacy task.
- The robot is already at a non-pharmacy station: mark a return request and wait for HMI continue.
- A return failure for one task does not block other tasks; that item returns
EMERGENCY_RETURN_REQUIRES_RECOVERY. - It returns
notifiedRobotsand the processing result for each task.emergency.changedis published after the Mongo pause succeeds.
HospitalMissionControlService injects only EmergencyOperationsPort and Hospital configuration. Repositories, Redis runtime state, HMI control, cancellation safety, and queue creation stay encapsulated in EmergencyOperationsService, so Hospital does not know how Operations data is manipulated.
6.2 Medicine-Cart Continue Orchestration
When the HMI continues a task, Product API calls MedicineCartTaskWorkflow. This Hospital workflow then calls Operations TaskProgressionService:
- For a normal task, Operations queues the next station and returns a
continuedoutcome. - When an unhandled
returnToErPharmacyRequestedAtexists, Operations returnsemergency-return-requiredwith a transport-neutral task context. - Hospital chooses the return station, invokes
EmergencyOperationsPort.cancelAndQueueReturn(), and maps the result to the unchanged HTTP response shape.
The Operations outcome contains no Hospital class, HTTP DTO, or CSH configuration:
type TaskContinuationOutcome =
| { kind: 'continued'; result: ContinueTaskResponse }
| { kind: 'emergency-return-required'; task: EmergencyContinuationContext };There is no Operations-to-Hospital reverse dependency and no Hospital provider token bound back into TaskModule. TaskProgressionService is registered normally by TaskModule; MedicineCartTaskWorkflow is registered by CshHospitalModule and exported to the controller.
6.3 Configuration Flow
Operations reads the following values through SiteOperationsConfig:
siteNamereturnHomeStationtaskSimulation
CshHospitalConfig implements this contract. AdaptersModule binds it to SITE_OPERATIONS_CONFIG using useExisting.
To add site-level configuration required by Operations:
- Add a transport-neutral field to
SiteOperationsConfig. - Implement it in the Hospital configuration while preserving backward-compatible defaults.
- Add configuration unit tests and Operations use-case tests.
- Update
.env.exampleand deployment documentation.
Do not inject ConfigService into Operations to read hospital-specific environment variables. Do not read process.env directly while loading application modules.
6.4 Business Event Publication
Operations and Hospital publish these events through OperationsEventPublisher:
robot.state.changedtask.state.changedfleet.state.changedemergency.changed
The current implementation is consumed by a WebSocket adapter. This publisher supports dashboard and event notifications; it is not a durable event bus. It must not replace MongoDB state changes or queue transactions that are required to succeed.
7. Emergency Return-to-Pharmacy Flow and Invariants
sequenceDiagram
participant CMP
participant HospitalAPI
participant HospitalWF as HospitalMissionControl
participant Ops as EmergencyOperationsPort
participant Mongo
participant HMI as RobotControlPort
participant Queue as Task Queue Worker
participant RMF as FleetCommandPort
participant ReturnOp as EmergencyReturns
CMP->>HospitalAPI: ReturnToErPharmacy
HospitalAPI->>HospitalWF: requestForActiveTasks()
HospitalWF->>Ops: pauseFleetTaskAcceptance()
Ops->>Mongo: persist pause timestamp + version
HospitalWF->>Ops: listActiveTaskContexts()
Ops->>Mongo: load active tasks and ownership
Ops->>Ops: read normalized Redis runtime state
HospitalWF->>Ops: broadcastEmergency(EMERGENCY_STARTED)
Ops->>HMI: notify every robot
loop each active task
alt robot is moving
HospitalWF->>Ops: cancelAndQueueReturn()
Ops->>ReturnOp: claim by original task id
Ops->>ReturnOp: phase=CANCELLATION_REQUESTED
Ops->>RMF: cancel booking
RMF-->>Ops: explicitly accepted
Ops->>Mongo: cancel original task locally
Ops->>ReturnOp: phase=ORIGINAL_CANCELED
Ops->>Mongo: create return task + queue intent
Ops->>ReturnOp: COMPLETED / RETURN_QUEUED
else arrived at non-pharmacy station
HospitalWF->>Ops: markReturnRequested()
Ops->>Mongo: persist return request + emergency version
end
end
HospitalWF-->>CMP: affectedTasks + notifiedRobots
Queue->>RMF: dispatch return task safelyChanges to this flow must preserve all of the following:
- A fleet pause blocks ordinary tasks and Robot Keeper auto-charging tasks.
- Manual cancellation remains rejected while the active robot is emergency-paused.
- Only the internal emergency-return workflow may use
allowDuringEmergency: true. - Emergency return must also use
requireFleetConfirmation: true. Local queue deletion, task-history cancellation, and return creation happen only after RMF explicitly accepts cancellation. - A rejected or unknown RMF cancellation must not be retried automatically and must not create a return task. The durable operation remains available for manual reconciliation.
- The emergency-return cancellation uses
notifyAmr: false, preventing a normalCANCELEDscreen from replacing the emergency HMI screen. EMERGENCY_STARTEDandEMERGENCY_CLEAREDnotify robots in every state, not only active robots.- One HMI notification failure cannot block the fleet workflow; notification uses fire-and-forget failure isolation.
notifiedRobotscontains robot names for which notification was attempted. It does not mean every HMI acknowledged delivery.- An emergency return may bypass an unbacked stale reservation, but not tracking that clearly belongs to another task.
- Mission continuation first captures the current emergency versions and clears only matching return requests and fleet pauses. A newer version created concurrently must remain active, and
EMERGENCY_CLEAREDmust not be broadcast.
7.1 Emergency Version Concurrency Rules
- Every
ReturnToErPharmacyevent creates a new version. Robots that are already paused keep their originaltaskAcceptancePausedAt, while their version advances to the newest event. MissionContinuescaptures all current versions before conditional clearing. An emergency that begins after the capture does not match and cannot be cleared by the older request.- Older rows without a version are treated as the
nulllegacy version so the first clear after upgrade can remove them safely. EMERGENCY_CLEAREDis published and broadcast only after confirming that no paused robots remain.
7.2 Emergency Return Idempotency and Recovery
The MongoDB EmergencyReturns collection uses the original taskHistoryId as its unique _id. Concurrent HMI continues or redelivered events for one task have exactly one operation owner. A replay after completion returns the same returnTaskId.
state | phase | Meaning and action |
|---|---|---|
PROCESSING | CLAIMED | Cancellation intent has not been recorded; inspect backend logs and Mongo state if it remains here |
PROCESSING / FAILED | CANCELLATION_REQUESTED | RMF cancellation may not have been sent, may have been rejected, or may be unknown; reconcile with RMF before any replay |
PROCESSING / FAILED | ORIGINAL_CANCELED | A legacy or not-yet-allocated flow; it cannot prove that no return was created before a crash, so only manual reconciliation is allowed |
PROCESSING / FAILED | RETURN_ALLOCATED | returnTaskId was persisted first; TaskHistory and TaskQueue can be completed idempotently with that same ID |
COMPLETED | RETURN_QUEUED | The return exists; use returnTaskId for subsequent tracking |
The API never automatically reclaims a FAILED operation. On-call staff must reconcile the RMF booking, original TaskHistory, robot ownership, return TaskHistory, and TaskQueue before following the site procedure to create a return manually or correct the operation. Deleting an EmergencyReturns record re-enables the side effect and is forbidden until reconciliation is complete.
Run yarn diagnose:emergency-returns to list operations that have remained unresolved for more than five minutes and view the operation, original task, known return task, queues, and robot ownership together. Use --stale-minutes 0 for every unresolved operation or --task-id <original-task-history-id> for one exact operation. This command is MongoDB read-only and always reports automaticReplayAllowed: false; recommendedAction classifies the reconciliation work and never authorizes an automatic recovery.
The new flow persists a preallocated returnTaskId in RETURN_ALLOCATED before it creates return TaskHistory. Only FAILED + RETURN_ALLOCATED is eligible for the guarded recovery command. The command uses that same ID to fill in a missing TaskHistory or TaskQueue and never replays RMF cancellation. Copy the exact updatedAt from a fresh diagnostic and provide the operator, reason, and confirmation flag. A conditional Mongo update prevents a stale diagnostic or concurrent operator from obtaining recovery ownership, and every result is appended to recoveryAudit.
8. Data Authority and Consistency
| Data | Authoritative source | Notes |
|---|---|---|
| Task history and task ownership | MongoDB | Redis or RMF events do not replace task history |
| Queue lifecycle, leases, and attempts | MongoDB | Claim and update operations must remain atomic |
taskAcceptancePausedAt and taskAcceptancePauseVersion | MongoDB | Fleet-wide emergency gate and conditional clear |
returnToErPharmacyRequestedAt and emergencyRequestVersion | MongoDB | Task-level emergency request waiting for HMI continue |
| Emergency return operation | MongoDB EmergencyReturns | Original-task idempotency, cancellation phase, and return task id |
| RMF status, battery, and location | Redis runtime state | Has a TTL and may be missing |
fullCharging and obstacle telemetry | Redis runtime state | Exposed as an overlay by GET robots |
| HMI busy lease | Redis | Must be included in availability decisions |
| RMF booking reconciliation | Mongo task/queue plus RMF inbound | Re-delivered inbound events must be idempotent |
Redis reads must distinguish between:
missing: no runtime state currently exists for the robot.failure: Redis could not be read.
Dispatch, availability, auto-charge, and emergency task evaluation must fail closed on failure. A read failure must never be interpreted as an idle, available, or empty state. A Mongo repository failure must not be mapped to “robot not found.”
9. Queue and Dispatch Safety Rules
The queue lifecycle is:
PENDING → PROCESSING → DISPATCHING → IN_PROGRESS → COMPLETED / FAILEDAll contributors must preserve these rules:
- A worker may process only the record it successfully claimed through an atomic MongoDB
findOneAndUpdate. - Claiming writes
workerId,leaseExpiresAt, andattemptCount. DISPATCHINGmust be persisted before calling RMF.- An explicit RMF rejection may be retried within the configured attempt limit.
- An ambiguous network result must not be dispatched automatically again, because RMF may already have created the task.
- Expired
PROCESSINGleases may be safely reclaimed; unreconciledDISPATCHINGrecords move to an explicit failed state. taskHistoryIdis the idempotency root.- A multi-station task reuses the same queue record and increments
dispatchSequence. - Robot Keeper must create charging tasks with an atomic deduplication condition.
- An ordinary task without docking map metadata must still be dispatchable. Charging metadata is required only when the task genuinely needs docking activity.
10. Dependency Injection and Port Binding
Inject a port through its Symbol token, not through a concrete adapter class:
constructor(
@Inject(FLEET_COMMAND_PORT)
private readonly fleetCommand: FleetCommandPort,
) {}Bind the adapter with useExisting in the adapter module:
RmfFleetAdapter,
{
provide: FLEET_COMMAND_PORT,
useExisting: RmfFleetAdapter,
}useExisting ensures that the token and concrete provider resolve to the same singleton instead of creating the adapter twice.
When adding a port:
- Place it in
application/portsof the feature that owns the requirement. - Define transport-neutral commands and results.
- Implement it in an adapter.
- Bind it with
useExistinginAdaptersModuleand export the token. - Test the application service with a fake port or mock.
Place a contract in operations/ports only when it is shared by different top-level modules and represents explicit cross-module collaboration.
11. Common Change Scenarios
11.1 Adding a General Task Rule
- Put the invariant in
operations/task/domainor the task application use case. - If robot or fleet data is required, depend on an existing contract; do not import a Mongo repository class.
- If RMF needs a new capability, extend
FleetCommandPortand the RMF adapter. - Add task service tests and, when dispatch is involved, queue and RMF adapter tests.
11.2 Adding a Hospital Event
- Validate the transport request in a Hospital DTO.
- Interpret the hospital event in
csh-hospital/workflows. - Use public Operations capabilities for task and robot actions.
- If Operations lacks a complete use case, add one instead of letting Hospital modify more repository fields directly.
- Preserve HTTP status codes, error codes, and response shapes unless an API version change has been formally agreed.
11.3 Adding Another Hospital
Do not copy Operations. Create another Hospital module, implement its site configuration and orchestration workflow, and reuse EmergencyOperationsPort plus feature public capabilities.
Before integrating another hospital, identify:
- Station naming and return-station rules.
- Emergency broadcast and HMI endpoint differences.
- Authentication and HIS/CMP payloads.
- Medication and employee data sources.
- Whether medicine-cart continue needs a different Hospital policy or workflow.
- How
AppModuleselects exactly one Hospital module and site configuration for a single-site deployment.
11.4 Adding or Replacing an External Integration
- Keep Axios, ROS, and Socket.IO clients in adapters.
- Inject typed configuration into adapters; do not read raw environment variables in application code.
- Map external failures to results or errors understood by the application layer.
- Normalize raw RMF payloads in the inbound adapter before calling
UpdateRobotStateService.
11.5 Changing an HTTP DTO or Route
- Keep DTOs in the API layer. Do not add Swagger or class-validator decorators to domain models.
- Keep application commands/results separate from HTTP DTOs.
- Preserve existing routes, status codes, error codes, and request/response shapes. If a breaking change is unavoidable, introduce a new API version and coordinate with consumer teams.
- Compare Swagger paths, methods, statuses, and schemas after the change.
12. Import Rules
Allowed
// A controller uses an application capability through the feature public entry.
import { TaskRequestService } from '../operations/task/public';
// An application service injects its own or another feature's contract.
import {
ROBOT_STATE_STORE,
RobotStateStore,
} from '../../robot/application/ports/robot-state-store.port';Forbidden
// A controller bypasses the public entry and imports the feature internals.
import { TaskRequestService } from '../operations/task/application/task-request.service';
// An application service imports a concrete adapter.
import { RmfFleetAdapter } from '../../../adapters/rmf/rmf-fleet.adapter';
// Domain code depends on NestJS, Mongoose, or Swagger.
import { Injectable } from '@nestjs/common';The current .eslintrc.js checks that:
- Domain code does not import frameworks or infrastructure.
- Application services, ports, and workflows do not import adapters, infrastructure, Product API, Hospital implementations, or concrete transport packages such as Axios, Mongoose, Redis, ROSLIB, or Socket.IO.
- Controllers do not import repositories, adapters, infrastructure, or feature-internal application trees.
test/service/import-boundaries.service.spec.ts uses intentionally invalid imports to verify the ESLint overrides. Run it whenever override order or globs change so a later Hospital override cannot silently replace controller restrictions.
Architecture still requires reviewer judgment. ESLint can check import patterns, but it cannot determine whether a service exposes too much internal capability.
13. Current Boundaries and Further Convergence
13.1 The Emergency Facade Is the Single Cross-Module Entry
HospitalMissionControlService no longer injects feature-local repository or control ports. It uses only behavior-oriented capabilities from EmergencyOperationsPort:
listActiveTaskContextsbroadcastEmergencypauseFleetTaskAcceptance/clearFleetTaskAcceptancemarkReturnRequested/clearReturnRequested/markReturnHandledcancelAndQueueReturn
EmergencyOperationsService implements this port inside Operations and coordinates task, robot, fleet, runtime state, HMI, and queue behavior. When adding emergency capability, expose a complete business operation instead of repository CRUD.
13.2 public/index.ts Is an Enforced Cross-Module Import Boundary
Hospital may use explicitly exported capabilities such as TaskProgressionService from operations/task/public, but it must not import operations/*/application/**. .eslintrc.js enforces this rule for src/csh-hospital/**/*.ts.
13.3 Task Progression and Hospital Orchestration Have Separate Registration
TaskProgressionService is a general task application capability registered by TaskModule. MedicineCartTaskWorkflow is CSH Hospital orchestration registered and exported by CshHospitalModule for Product API controllers.
Do not move Hospital policy back into TaskProgressionService, and do not make Operations depend on CshHospitalModule. New branch outcomes should remain transport-neutral and be interpreted by the Hospital workflow.
13.4 Each Deployment Currently Uses One Hospital Policy Set
The current composition expects each deployment to select one Hospital module and one SiteOperationsConfig. Supporting several hospitals in one process requires site/hospital routing; one global configuration token is insufficient.
14. Testing and Acceptance
14.1 Minimum Validation Matrix
| Change scope | Required validation |
|---|---|
| Pure domain/value object | Lint, build, feature unit tests |
| Application service | Above plus service tests |
| Module/provider/public entry | Above plus module-boundary tests |
| Port/adapter | Above plus adapter contract tests |
| Queue/emergency/runtime state | MongoDB/Redis integration tests plus E2E |
| HTTP DTO/route | E2E plus Swagger contract comparison |
| Deployment/configuration | Production image build plus startup smoke test |
14.2 Standard Commands
Use Node.js 24.13.0 to match the pipeline:
yarn install --frozen-lockfile
yarn lint:check
yarn test:architecture
yarn build
yarn test test/service --runInBand
yarn test:integration --runInBand
yarn test:e2e --runInBandService integration tests and E2E require MongoDB, Redis, and a test .env. Connection failures in a bare checkout without these dependencies are neither program regressions nor successful validation.
The Jest configurations intentionally do not use forceExit. A process that prints test results but does not terminate is still a lifecycle failure. E2E suites must close the Nest application through AppHelper.closeAgent(), test databases through MongoHelper.close() so the underlying pool is closed, and new long-running adapters or workers must abort or drain work in a Nest shutdown hook.
Docker Compose validation is recommended:
docker compose up -d mongodb redis
docker compose run --rm backend yarn test test/service --runInBand
docker compose run --rm backend yarn test:e2e --runInBandThe effective Compose configuration must point the test container to the mongodb and redis service names, not to localhost inside the container.
14.3 Critical Regression Cases
Any Operations/Hospital change should verify at least:
- An ordinary single-station task without docking metadata can still be dispatched.
- Only one worker wins a multi-worker claim race.
- Dispatch timeout or an unknown result does not trigger redispatch.
- Multi-station continuation increments the sequence correctly.
- Duplicate RMF terminal events do not complete a task twice.
- Redis failure causes dispatch and auto-charge to fail closed.
- Redis or Mongo read failure makes emergency evaluation fail closed while the fleet pause remains active.
- Emergency pause blocks normal tasks and Robot Keeper.
- An older
MissionContinuescannot clear a newer emergency version. - Concurrent emergency returns for one original task produce only one cancel/create side effect.
- RMF cancel rejection or timeout does not delete the local queue or create a return; the operation phase remains available for reconciliation.
- Emergency return follows the stale-reservation rules.
- Manual cancellation remains rejected during emergency pause.
- Fleet-wide HMI broadcast includes robots in every state and preserves
notifiedRobots. - Obstacle and full-charging telemetry overlays do not regress.
- Docking-arrival decision-manager activity and activity order remain unchanged.
15. Pull Request Collaboration Process
15.1 Author Checklist
- The owning feature or Hospital boundary has been identified.
- No hospital name or hospital-specific environment condition was added to Operations.
- Controllers perform only validation, mapping, application calls, and response mapping.
- Application code does not import a concrete adapter.
- New ports use string IDs and transport-neutral types.
- Raw RMF, Axios, Mongoose, or Redis representations do not cross adapter boundaries.
- Queue, emergency, idempotency, and fail-closed invariants remain intact.
- Relevant unit, integration, and E2E tests pass.
- Documentation and
.env.exampleare updated for API or configuration changes.
15.2 Reviewer Checklist
- Business rules are placed by ownership, not by the source of the request.
- Any newly public service genuinely needs to be exported from the feature’s
publicentry. - Hospital has not gained another repository-port dependency; if it has, consider an Operations use case instead.
- Operations has not gained a dependency on a concrete CSH/CMP/HIS implementation.
- Error mapping and HTTP contracts remain compatible with existing consumers.
- Failure and ambiguous outcomes are distinguished correctly.
- Fire-and-forget is used only for notifications that permit partial failure.
- Data authority and transaction/atomic-update behavior are explicit.
15.3 Merge and Deployment
- Each commit should build and be independently revertible.
- Keep data migrations separate from code-only refactors.
- Run queue schema migrations according to
task-queue-v2-migration-runbook.md: stop the old worker, migrate, and only then start the new version. - After starting the production image, follow
deployment-runbook.mdto validate MongoDB, Redis, RMF, Fleet Manager, HMI, and Hospital integrations.
16. Suggested Team Ownership
| Area | Primary owner | Required collaborators |
|---|---|---|
| Operations domain/application | Backend/Operations team | RMF, Hospital, QA |
| CSH Hospital workflows/integrations | Hospital Integration team | Backend, Hospital API owner, QA |
| RMF/Robot adapters | Robot Integration team | Backend, fleet vendor |
| Product/Hospital HTTP contract | API owner | Consumer teams, QA |
| MongoDB/Redis schemas and migrations | Backend/Data owner | SRE, QA |
| Deployment/configuration | SRE/Platform | Backend, Hospital IT |
Before starting cross-team work, identify at least the owner, API/port contract, data authority, failure semantics, idempotency strategy, test environment, and deployment order.
17. Important File Index
| File | Purpose |
|---|---|
src/operations/operations.module.ts | Operations feature composition |
src/operations/*/*.module.ts | Feature providers and exports |
src/operations/*/public/index.ts | Feature public import entries |
src/operations/ports/ | Cross-top-level-module contracts |
src/operations/workflows/ | Emergency facade, Queue Worker, and Robot Keeper |
src/csh-hospital/csh-hospital.module.ts | Hospital providers and workflow composition |
src/csh-hospital/workflows/hospital-mission-control.service.ts | Emergency and return-to-pharmacy workflow |
src/csh-hospital/workflows/medicine-cart-task.workflow.ts | HMI task progression and emergency-return orchestration |
src/adapters/adapters.module.ts | Composition of ports and concrete adapters |
.eslintrc.js | Static import-boundary rules |
test/service/module-boundaries.service.spec.ts | Module-structure and provider-graph validation |
docs/modular-monolith-architecture.md | System architecture summary |
docs/deployment-runbook.md | Build, deployment, and rollback instructions |
docs/task-queue-v2-migration-runbook.md | Queue migration procedure |
18. Glossary
- Feature public entry: A feature’s
public/index.ts, which explicitly lists externally usable application services. - Port: An abstract capability required by application code and injected through a NestJS token.
- Adapter: A technical implementation of a port, such as a MongoDB repository or RMF client.
- Workflow: A business process coordinating several use cases or features.
- Runtime state: Current RMF/robot state stored primarily in Redis.
- Task ownership: The task currently assigned to a robot, with MongoDB as the authority.
- Fail closed: Reject dispatch or automation when external state cannot be verified instead of assuming it is safe.
- Dispatch reconciliation: Determining the actual dispatch result from inbound RMF booking/status data after an ambiguous response.
- Idempotency root: The stable identifier used to recognize one business task; currently
taskHistoryId.