An RTOS makes concurrency easier to express and easier to get wrong. The difference between a stable product and a Heisenbug factory is usually not the kernel — it is whether tasks have explicit budgets.
Name the rates
List every periodic activity with a period, deadline, and worst-case execution time estimate:
| Task | Period | Deadline | Notes |
|---|---|---|---|
| Control | 1 ms | 1 ms | no blocking calls |
| Comms | 10 ms | 20 ms | may wait on TX |
| Logging | 100 ms | soft | lowest priority |
If you cannot fill the table, you do not have a design yet — you have a hope.
Priorities follow deadlines
Rate-monotonic thinking is a good default for simple systems: shorter period → higher priority. Then audit blocking: a high-priority task that waits on a mutex held by a low-priority task is a priority inversion waiting to happen. Use priority inheritance or, better, redesign the shared resource.
Stacks are measured, not guessed
Start with generous stacks, run watermark hooks, and shrink with evidence. Print high-water marks in bring-up builds:
UBaseType_t hw = uxTaskGetStackHighWaterMark(NULL);On Zephyr, thread analyzer tooling plays a similar role. Blind 2048 byte stacks everywhere waste RAM and still overflow the one task that JSON-encodes logs.
Leave margin for the field
Lab WCET is not field WCET. Interrupts, cache effects, and flash wait states expand execution time. If your 1 kHz loop uses 80% of its period on the bench, it will miss in production.
Budgets turn “the RTOS felt fine” into engineering you can defend.