Resolving Bus Error on IOT2050 Advanced After eMMC Install

David Krause15 min read
Other TopicSiemensTroubleshooting
Licensed PE Working through this on a live machine? A Maine-licensed engineer can take it from here — included with IMD hardware, by the hour for everything else. Book an engineer

1. Problem Overview

The SIMATIC IOT2050 Advanced (order suffix ...-1YA2) is shipped with an internal eMMC that can host the operating system and runtime services, including the Node-RED flow engine that ships preconfigured on the official Example Image. When a user formats the eMMC manually with mkfs.ext4 /dev/mmcblk1 and then flashes the Example Image via the documented SD-card-boot + USB-transfer procedure, the device usually comes back up cleanly. In a small number of deployments, however, the gateway reaches a partially working state in which the IP stack responds, the root login and password change succeed, and the network ports are configured correctly, but Node-RED refuses to start. The symptom reported is a bare Bus error returned by the shell when the user runs the node-red command from a Putty/SSH session, and an empty TCP socket on port 1880 because the service never reaches a listening state.

The message Bus error is not a Node-RED diagnostic, a Node.js error, or a Linux kernel panic. It is the shell's translation of a process termination caused by the POSIX signal SIGBUS (signal number 7 on Linux). The shell prints Bus error the same way it prints Segmentation fault for SIGSEGV, and the only useful information it carries is the fact that the process died before it could produce any application-level output. Recovering the gateway requires identifying what triggered SIGBUS during startup and re-establishing a clean, verifiable copy of the Example Image on the eMMC.

Field note: A Bus error that appears only for one or two specific binaries (node-red, npm, node) while other commands run fine is almost always related to a corrupted executable, a corrupted shared library the executable maps into memory, or an eMMC read failure while the kernel is satisfying an mmap-backed page fault. Treat it as a storage-integrity problem before treating it as a software problem.

2. Understanding the Bus Error Signal

On a Linux system, the kernel raises SIGBUS when a process performs an invalid access to a memory region that is not aligned to the access width, or when the backing storage that should satisfy the access is unavailable, truncated, or physically defective. The most common triggers in industrial-embedded deployments are:

Trigger Mechanism Typical indicator
Unaligned access to mmapped file on strict-alignment architecture The load address or file offset is not a multiple of the access size and the kernel cannot emulate the access. Sporadic, depends on ASLR and file layout.
Filesystem corruption on the executing partition mmap()ed pages of the ELF or shared library cannot be read because the inode is bad, the block is bad, or the file is truncated. Repro on every invocation; other binaries in the same path may also fail.
eMMC read error on a mmap-backed block The kernel pages data in from the eMMC device and the controller returns a CRC error or timeout on the read. Correlates with mmc0/mmc1 errors in dmesg.
Truncated or partially written executable An image-write operation (dd, bmaptool, balenaetcher, bzip2 | dd) was interrupted, leaving the file shorter than the loader expects. File size in ls -l differs from the source image's payload size.
Bad ELF interpreter (/lib/ld-linux-*.so.*) The dynamic loader itself is corrupt; every dynamically linked binary fails to start. All dynamically linked binaries (including ls, cat from busybox) throw Bus error.

The IOT2050 Advanced platform uses an ARM SoC (Texas Instruments Sitara family) on which the kernel is configured for alignment-tolerant user-space access in most distributions. This means that, in practice, the most likely cause on this hardware is storage-side, not alignment-side. Specifically, the problem is almost always either (a) an eMMC block failure while the kernel pages in a node.js V8 mapping or (b) a partially written binary left behind by an interrupted flash operation.

3. Root Cause Analysis for the Reported Failure

The deployment path that produces the failure combines three risk factors that interact badly:

  1. Pre-format of the eMMC with mkfs.ext4. The IOT2050 factory default uses a different partition layout than the standalone mkfs.ext4 /dev/mmcblk1 would create. Formatting the whole block device with a single ext4 filesystem destroys any partition table the factory layout depends on and can interfere with the image-write tool's expectation of an empty block device.
  2. Two-stage install (SD card boot + USB transfer). When the Example Image is transferred to the eMMC by a USB key on a target that is booted from SD, the image writer reads the WIC image, streams the partitions to /dev/mmcblk1, and then synchronizes. If the USB key is slow, the target loses power, or the IOT2050 is rebooted before the write finishes, the userland (rootfs) is left in a state where the files are present but the contents of the mmap()ed region are wrong.
  3. Node.js uses mmap() heavily. The V8 engine mmaps its heap snapshot, and the Node.js binary itself relies on shared libraries. If the underlying block is bad, the page fault handler returns -EIO to the user process and the kernel delivers SIGBUS.

The combination of these three factors explains why the gateway can come up far enough to accept a password change and configure a port, but not far enough to start Node-RED: the early userspace (init, login, ifconfig, dropbear/sshd) is built from static or simple dynamic libraries that fit into healthy page clusters, while the Node.js binary hits a corrupt page that was clobbered during the partial write. This pattern - "everything works except one or two specific binaries" - is the field signature of a partial-flash / eMMC-block-corruption problem on the IOT2050.

4. Pre-Fix Diagnostic Procedure

Before re-imaging the device, capture the state so that the failure can be confirmed and a clean baseline established. Connect to the gateway over the serial console or over SSH as root and run the following commands. Save the full output to a file on the USB key or to an SCP target.

  1. Confirm the kernel and image versions.
    uname -a
    cat /etc/os-release
    cat /etc/version
    These should report the Example Image V1.1.1 build expected on the IOT2050 Advanced.
  2. Reproduce the failure deterministically.
    node --version
    npm --version
    which node-red
    ls -l $(which node-red)
    node-red
    The first three commands isolate whether the failure is in the node binary, the npm shebang, or the node-red wrapper. In the reported case, node --version ran but npm --version and node-red did not. That is consistent with a corrupt page inside the npm/node-red binary, not the node interpreter.
  3. Check the kernel ring buffer for eMMC errors.
    dmesg | grep -iE 'mmc|emmc|i/o error|sigbus|bus error|EXT4-fs'
    Lines containing I/O error, mmcblk1: error, read error, or EXT4-fs error confirm a storage problem and remove any doubt about software causes.
  4. Confirm the on-disk block device and filesystem state.
    lsblk -f
    mount | grep mmcblk1
    blkid /dev/mmcblk1*
    The lsblk -f output should show a small boot partition (mmcblk1p1, vfat) and a root partition (mmcblk1p2, ext4) on the eMMC. If only a single mmcblk1 with one big ext4 filesystem is present, the pre-format step has wiped the expected partition layout, which is itself a strong indicator of an unaligned install.
  5. Check the integrity of the node binaries on disk.
    md5sum /usr/bin/node /usr/bin/node-red /usr/bin/npm \
           /usr/lib/node_modules/node-red/red.js 2>/dev/null
    On a clean install these values are stable for a given image build. Compare them with the values from a second IOT2050 Advanced that has the same Example Image V1.1.1 on it. Any difference confirms a partial flash.
Important: Do not skip step 3. The dmesg output is the single piece of evidence that distinguishes a recoverable software issue (re-flash the eMMC) from a hardware issue (replace the eMMC / return the unit for RMA). If dmesg shows persistent eMMC read errors after a clean re-flash, escalate before reinstalling a third time.

5. Solution: Clean Reinstallation of Example Image V1.1.1

The known-good recovery on the IOT2050 Advanced is a clean re-flash of the Example Image. The procedure below mirrors the official Setting up the SIMATIC IOT2050 reference workflow used by the SIMATICmeetsLinux project.

5.1 Prerequisites

  • Example Image V1.1.1 (WIC image) downloaded from Siemens Industry Online Support, SHA-256 verified against the value published with the release.
  • SD card (>= 8 GB, class 10 or better) flashed with the boot portion of the image.
  • USB stick formatted as ext4 or FAT32, holding the full WIC image and the image-write tool used by the procedure.
  • Uninterruptible power, or at minimum an isolated power circuit that cannot be disturbed during the write.
  • Serial console (115200 8N1) on the IOT2050 Advanced, or a configured SSH path to the gateway so that you can re-flash headlessly.

5.2 Reset the eMMC to a Known State

Boot the IOT2050 Advanced from the SD card. Do not boot from the eMMC. Once the SD card image is running, identify the eMMC block device and wipe it cleanly so that no stale partition table interferes with the write:

lsblk
# Confirm the eMMC is /dev/mmcblk1 (not /dev/mmcblk0, which is the SD card)
ls -l /dev/mmcblk*

# Wipe partition table and header areas
dd if=/dev/zero of=/dev/mmcblk1 bs=1M count=16 conv=fsync
# Force a fresh partition table re-read
partprobe /dev/mmcblk1 || true
blockdev --rereadpt /dev/mmcblk1 || true
Why this step matters: A pre-existing mkfs.ext4 /dev/mmcblk1 install (the same step that preceded the failure in the source report) leaves a single ext4 superblock but no GPT/MBR partition table. When the WIC writer later writes the boot (p1) and root (p2) partitions of the Example Image, it expects to be writing into a clean device. A stale superblock at byte offset 1024 of the raw device can confuse the writer and lead to truncated outputs, which is the immediate precursor to a Bus error on a mapped executable.

5.3 Write the Example Image to the eMMC

Use the documented image-write procedure, which on the IOT2050 is the WIC-based installer run from the booted SD card. The reference workflow lives in the SIMATICmeetsLinux Setting up the IOT2050 guide and the Siemens Industry Online Support download bundle. The high-level steps are:

  1. Insert the USB stick carrying the WIC image and the image-write tool.
  2. Mount the USB stick (e.g. mount /dev/sda1 /mnt).
  3. Locate the WIC image (e.g. example-image-iot2050.wic or *.wic.zst) and verify its SHA-256 against the Siemens-published value.
  4. Run the WIC writer to stream the image to /dev/mmcblk1. Wait for the writer to report success. Do not interrupt it.
  5. Power down the IOT2050 Advanced, remove the SD card, and power it back on so that it boots from the eMMC.

5.4 First-Boot Configuration

After the first boot from the eMMC:

  1. Log in as root over the serial console and immediately change the root password.
  2. Configure the network interface as documented (e.g. /etc/network/interfaces.d/ on the Example Image).
  3. Confirm that node-red is enabled and starts at boot:
    systemctl status nodered
    systemctl enable --now nodered
  4. Open a browser and navigate to http://<iIP2050 IP>:1880. The Node-RED editor must load within a few seconds of the service starting.

6. Service Verification Checklist

Before declaring the gateway healthy, run the full verification matrix below. Every line must return the expected state; if any line fails, do not close the change ticket - re-flash and re-run.

Check Command Expected
Image identity cat /etc/version Reports Example Image V1.1.1 build identifier.
Filesystem layout lsblk -f Two partitions on mmcblk1: p1 vfat (boot), p2 ext4 (root).
Kernel clean dmesg | grep -iE 'error|fail' No eMMC or EXT4 error lines.
Node binary OK node --version Returns a version, no Bus error.
npm OK npm --version Returns a version, no Bus error.
Node-RED binary OK node-red --help Returns usage text, no Bus error.
Service active systemctl is-active nodered active.
Service enabled systemctl is-enabled nodered enabled.
Port 1880 listening ss -tlnp | grep 1880 Node.js process owns the socket.
Editor reachable curl -fsS http://127.0.0.1:1880/ HTTP 200 with Node-RED HTML.
MD5 of node binaries md5sum /usr/bin/node /usr/bin/node-red /usr/bin/npm Matches the values from a known-good unit of the same image build.

7. eMMC Health Verification (When the Failure Recurs)

If the same Bus error returns after a clean re-flash, the eMMC itself is the prime suspect. The IOT2050 Advanced eMMC can be read-tested in place from the booted SD-card image without unmounting it. This avoids the cost and downtime of an RMA when the unit is otherwise healthy.

  1. Read the eMMC device identification.
    cat /sys/block/mmcblk1/device/serial
    cat /sys/block/mmcblk1/device/manfid
    cat /sys/block/mmcblk1/device/oemid
    Record the manufacturer and serial for the service log.
  2. Read the eMMC health counters (extended CSA / device health).
    mmc extcsd read /dev/mmcblk1 | grep -E 'DEVICE_LIFE_TIME_EST|DEVICE_LIFE_TIME_EST_TYP|PRE_EOL' || true
    DEVICE_LIFE_TIME_EST_TYP is the JEDEC-defined wear indicator. A value of 0x01 means the device has reached a small percentage of its rated endurance; 0x0A and above usually mean the device is near end-of-life. If the value is already at 0x0A on a freshly deployed unit, escalate to RMA.
  3. Run a non-destructive read test.
    badblocks -nsv -b 4096 /dev/mmcblk1
    badblocks -n performs a read-only pass. Any error here confirms a defective block and is grounds for replacement.
  4. Check for unreadable files.
    find /mnt/rootfs -type f -exec dd if={} of=/dev/null bs=4096 status=none \; 2>&1 | grep -v '^0'
    Any line other than 0+0 records in indicates a file that cannot be read back. This is the most direct way to find the exact path that is causing the Bus error at process start.
Wear-leveling note: The IOT2050 Advanced eMMC has internal wear leveling, so a few bad blocks are normal over the life of the device. What is not normal is a bad block inside an active, mounted rootfs partition with an active, running daemon. That combination is what produces the Bus error on the next access. A single pass of badblocks is enough to make the call; do not run multiple destructive passes on a production device.

8. Node-RED Service Recovery on an Already-Flashed Device

If a re-flash is not an option (for example, the device is in production with custom flows already deployed), it is sometimes possible to recover Node-RED without re-imaging, by replacing only the corrupt binaries. The procedure below is riskier than a clean re-flash and should be the second line of defense, not the first.

  1. Stop the Node-RED service:
    systemctl stop nodered
  2. Back up the user's Node-RED home directory before touching anything:
    cp -a /root/.node-red /root/.node-red.bak.$(date +%s)
  3. From a second, known-good IOT2050 Advanced (same image build), or from the original WIC image mounted read-only, copy the affected files into place. Common candidates are:
    cp -a /usr/bin/node /usr/bin/node-red /usr/bin/npm /usr/bin/
    ldconfig
  4. Re-run the verification matrix from Section 6. If the binaries still produce Bus error, the corruption is in the shared libraries or in the on-disk inode metadata and a clean re-flash is the only safe path.
  5. Restart the service and confirm that the editor loads:
    systemctl start nodered
    journalctl -u nodered -n 50 --no-pager

9. Prevention and Best Practices

The reported failure is preventable. Apply the following field practices on every IOT2050 Advanced deployment to keep Node-RED and the rest of the Example Image serviceable.

  • Do not pre-format the eMMC with mkfs.ext4 before flashing the Example Image. The WIC writer in the official Setting up the IOT2050 flow expects a clean block device. Wipe the device with a zero-write (dd if=/dev/zero of=/dev/mmcblk1 bs=1M count=16) only.
  • Verify the WIC image hash before writing. sha256sum -c image.sha256 is mandatory. A truncated or corrupted WIC produces the same partial-flash signature as a power loss during write.
  • Power the IOT2050 from a stable source during image writes. Industrial cabinets with weak 24 V rails and long DC runs are the most common source of mid-write power loss.
  • Always run the verification matrix from Section 6 after a flash. A gateway that boots and accepts a password change is not a gateway that is healthy; only the full matrix can prove that.
  • Schedule periodic eMMC health checks. Read DEVICE_LIFE_TIME_EST_TYP from mmc extcsd read /dev/mmcblk1 on a maintenance window once per quarter. Plan a replacement before the value reaches 0x05.
  • Back up the Node-RED flows file regularly. The flows are stored under /root/.node-red/flows_*.json by default. Export them to the same backup target as the PLC project so that a re-flash is a five-minute job, not an all-day job.
  • Keep the Example Image build pinned. Mixing V1.1.1 binaries with flows that target a different release is a common source of "works on my bench, fails in the cabinet" reports. Document the build string in the cabinet label.

10. Frequently Asked Questions

What does "Bus error" mean when Node-RED fails to start on the IOT2050 Advanced?

It is the shell's translation of SIGBUS (signal 7). On the IOT2050, the most common cause is a corrupted executable, a corrupted shared library, or a bad block on the eMMC (/dev/mmcblk1) that the kernel cannot page in. The fix is a clean re-flash of Example Image V1.1.1, not a Node-RED configuration change.

Why does only node-red and npm fail with Bus error while node --version works?

Node-RED and npm sit on a slightly different set of mapped pages than the node interpreter. A single bad block in the region of the filesystem that holds those binaries is enough to make the first access return SIGBUS while the node binary, sitting in a different cluster, still works. The selective failure is the field signature of partial-flash / eMMC-block corruption.

Can I keep the gateway in production by reinstalling only the Node-RED binaries?

Sometimes, yes. Stop the nodered service, back up /root/.node-red/, and copy node, node-red, and npm from a known-good IOT2050 Advanced (same image build) into /usr/bin/. Re-run the verification matrix. If the binaries still fail, the corruption is in shared libraries or in inode metadata, and a clean eMMC re-flash is required.

How do I tell if the eMMC is failing versus the image being bad?

Capture dmesg | grep -iE 'mmc|i/o error' after a clean re-flash. Repeated mmcblk1 read errors or a non-zero badblocks -nsv output confirm a hardware problem. In that case, escalate to RMA rather than re-flashing a third time.

Is the official Setting up the IOT2050 guide enough to recover a gateway that has this Bus error?

Yes, when followed end-to-end. The guide at SIMATICmeetsLinux Setting up the SIMATIC IOT2050 documents the SD-card-boot + USB-transfer procedure that produced the failure in the first place; running it once more, this time without the pre-format step and with a verified WIC hash, returns the gateway to a known-good state.

Back to blog