Discussion & Conclusion

What frequency shaping achieves

The weighted encoder changes the character distribution of an encrypted payload without replacing encryption. Uniform ciphertext bits are represented by letters emitted through a target-derived Huffman tree. The measured hostname entropy is lower than for regex/FTE and Base32.

This addresses a specific part of DNS tunnel detection: the difference between conventional encoded payloads and the biased letter distributions of ordinary names. It does not remove the fact that large amounts of information are being carried in DNS queries.

The distinction matters. A reduction in an individual indicator can be useful even when other signals remain, but it is not equivalent to complete concealment of a tunnel.

Throughput and transmission cost

Encoding capacity and per-character entropy are connected. Base32 allocates 5 bits to each character. Uniform lowercase output has a lower theoretical capacity of approximately 4.7 bits. The weighted English-frequency tree carries roughly 4.19 encrypted bits per emitted character on average.

The same file therefore requires a longer representation under weighted encoding. Given the DNS name limit, it must be divided across more queries. The benchmark confirms this trade-off: weighted transfers have the lowest measured entropy and the largest packet count, while Base32 has the highest throughput.

Reducing payload indicators can consequently strengthen a traffic-level indicator by increasing the volume needed to transfer the file.

Remaining traffic-level indicators

FTExfil still produces:

  • Long query names and labels carrying encoded chunks.
  • A high volume of DNS requests for substantial file transfers.
  • Many unique names under the same controlled domain.
  • Repeated communication with the receiving authoritative server.

These signals survive changes to the character distribution. Delaying queries changes their pace but does not eliminate the number of messages needed to carry the file.

Character distributions are not language

The emitted names remain unintelligible. They may resemble the letter frequencies of English or DNS more closely than uniformly chosen letters, but they do not reproduce common words or the usual dependencies between adjacent characters.

Unigram entropy measures the uncertainty of one character without its surrounding context. Bigram and trigram patterns describe sequences of two or three characters. Matching the first does not guarantee matching the others.

A name can therefore have a plausible overall entropy while still containing unnatural sequences. More contextual models are needed to explore this distinction.

A new statistical fingerprint

Huffman shaping introduces its own structure. For a codeword of length ℓ\ell, the expected emission probability is 2−ℓ2^{-\ell}. Over a sufficiently large sample, the output approaches those tree-defined probabilities rather than an arbitrary smooth target distribution.

This can create a recognizable fingerprint of the encoding. Whether it is sufficiently distinct to support reliable blocking depends on the detector and on overlap with legitimate traffic. That question is not answered by the entropy benchmark alone.

Variable-length capacity

The weighted encoder’s output length varies with the encrypted input. The English-frequency tree has codes as short as three bits and as long as nine. A sequence dominated by shorter codes emits more letters per input byte, while longer codes consume more bits before emitting a letter.

Using an average entropy of about 4.1 bits per character across 238 variable characters suggests approximately 976 bits, or 122 bytes, of raw representational capacity. But this is an average estimate, not a safe payload limit. Metadata and encryption reduce the usable portion, and a particular bitstream can expand beyond the expected length.

The conservative 80-byte benchmark chunk size and its encoding trials address this practical constraint for the chosen domain. Different suffixes, codebooks, and chunk lengths require renewed capacity checks.

Limits of the evaluation

The experiments use a controlled container network, a 1 MiB file, fixed query pacing, and five repetitions per configuration. They establish the relative performance and measured output entropy of those configurations.

They do not measure detector false-negative rates, behaviour across arbitrary recursive-resolver paths, or a complete model of real DNS traffic. Testing against operational DNS tunneling detection remains necessary before making stronger claims about evasion.

Future work

Contextual character models

More accurate DNS models could extend the current unigram approach. Two proposed directions are:

  1. Treat frequent bigrams or trigrams as symbols in a Huffman tree.
  2. Retain letter-level symbols but select the next Huffman tree according to previously emitted characters, using conditional frequencies.

Both aim to model relationships between characters rather than only their overall occurrence counts.

Evaluation against detectors

Further work could evaluate the output against real detection systems, including configurations built around Zeek, Snort, or Suricata. Such experiments would need to measure detection behaviour alongside throughput and packet volume, rather than infer it from hostname entropy.

Broader transport applications

Potential extensions include arbitrary protocol tunneling, a Tor pluggable transport, and frequency shaping for other encrypted protocols. These would broaden the current file-transfer pipeline and remain future research directions.

Conclusion

FTExfil demonstrates a functioning DNS file-transfer pipeline with configurable representations and optional ECDH key exchange. It can chunk a file, protect each chunk, encode the data into DNS query names, and reconstruct the original bytes at the receiver.

The novel use of Huffman coding reduces measured per-character entropy by shaping the visible distribution of encrypted data. This comes at the cost of throughput and additional DNS traffic. Regex-based FTE provides an intermediate trade-off, while Base32 remains the most efficient of the evaluated configurations.

The result is a proof of concept for encoding-based reduction of selected payload indicators, not a complete solution to DNS tunnel detection. The unresolved traffic signals, output-length variability, and lack of contextual language structure define the next research questions.