Mirth Connect encoding and character set issues come from a source system sending non-UTF-8 data, a connector charset setting that doesn't match the sender, HL7 escape sequences left undecoded, or a File Reader/Writer connector configured with the wrong file encoding. Fix it by setting the connector's charset to match the sender, explicitly configuring file encoding rather than relying on the platform default, handling HL7 escape sequences in the transformer, and converting to UTF-8 at the source where possible. Prevent recurrence by documenting each partner's encoding and testing with real special characters before go-live, since plain ASCII test data never reveals a mismatch.
Quick answer
Mirth Connect encoding and character set issues show up as garbled accented characters, replacement boxes in place of real text, or a java.nio.charset.MalformedInputExceptionin the server log when a connector can't decode incoming bytes with its configured charset. The message often still processes and sends successfully, which is what makes this quietly dangerous — the data itself is wrong, not missing.
Below is what causes the mismatch, the fix for each connector type, and how to catch corrupted characters before they reach a downstream system. Our free Mirth Health Check can check your connector's charset settings against the sender — part of the Mirth Connect support work we do for US healthcare teams.
What Causes Encoding and Character Set Issues?
An encoding mismatch happens whenever the bytes a system sends were written in one character set but read back using a different one, and healthcare messaging is especially exposed to this because patient names and addresses routinely contain accented or non-ASCII characters.
Source System Sending Non-UTF-8 Encoded Data
Many legacy systems still send data encoded in ISO-8859-1 or a Windows-specific codepage rather than UTF-8, and if the receiving connector assumes UTF-8 by default, any character outside the basic ASCII range decodes incorrectly.
Connector Charset Setting Mismatched With the Sender
Mirth's connectors let you configure an expected charset explicitly, and if that setting doesn't match what the sending system actually uses, every message will decode incorrectly in the same predictable way until it's corrected.
HL7 Escape Sequences Not Decoded Correctly
HL7 messages use their own escape sequences for special characters within fields, and if a transformer doesn't decode these correctly, the raw escape codes appear in the processed message instead of the intended character.
File Encoding Mismatch on File Reader and Writer Connectors
A File Reader or Writer connector configured with the wrong encoding will misinterpret or mis-write file contents in exactly the same way as a network connector, but the corruption is easier to miss since no connection error occurs at all.
How to Fix Encoding and Character Set Issues
Identify the sending system's actual encoding first, rather than guessing at random, since setting the wrong charset on the Mirth side will simply shift which characters look wrong instead of fixing the underlying mismatch entirely.
Set the Connector's Charset to Match the Sender
Configure the connector's charset setting explicitly to match what the sending system actually uses, rather than leaving it on a default that may not match every partner sending data into that same channel.
Explicitly Configure File Reader and Writer Encoding
Set the encoding on File Reader and Writer connectors directly in their connector settings rather than relying on the server's platform default, which can vary silently between operating systems and installations.
Handle HL7 Escape Sequences in the Transformer
Use Mirth's built-in HL7 escape sequence handling, or add explicit decoding logic in the transformer, so that escape codes for special characters are converted to their actual intended characters before the message reaches its destination.
Convert Encoding at the Source When Possible
Where you have influence over the sending system, standardizing on UTF-8 at the source eliminates the mismatch entirely rather than requiring Mirth to compensate for a legacy encoding on every single message received.
How to Prevent Encoding Issues From Recurring
Encoding problems tend to hide in plain sight because the message still processes successfully, so prevention depends on deliberately testing with the kind of characters that actually reveal a mismatch rather than assuming plain ASCII test data is sufficient on its own.
Document Each Integration Partner's Encoding
Keep a record of the character set each sending or receiving system actually uses, confirmed directly rather than assumed, so a new channel connecting to that same partner starts with the correct setting from day one.
Test With Real Special Characters Before Go-Live
Send test messages containing accented characters, apostrophes, and other non-ASCII content specific to your patient population before going live, since plain ASCII test data will never reveal an encoding mismatch that's actually present.
Standardize on UTF-8 Wherever You Control Both Ends
For integrations entirely within your own infrastructure, standardizing on UTF-8 everywhere removes an entire category of encoding mismatches that only exist because of inconsistent defaults across different systems and platforms.
Monitor for Replacement Characters in Processed Messages
Set up a simple check for the Unicode replacement character or common garbled patterns in processed message content, which can catch an encoding issue automatically rather than waiting for someone to notice a name displayed incorrectly.
When to call for help
If characters are still garbled after matching the charset setting, the cause is usually an undecoded HL7 escape sequence or a File Writer encoding that wasn't explicitly set. Our free Mirth Connect health checkchecks your connector's charset settings against the sender as part of the standard diagnostic.
Have a sample of the garbled output? Send us the log, or check pricing for ongoing support plans.
Seeing garbled characters right now? Free Mirth health check — a written 12-point audit report in 48 hours, no cost.