Efficient Data Storage in DNA Sequences Aided by p-gram Huffman Coding

Researchers from Yangzhou University have proposed an efficient approach to DNA data encoding that leverages p-gram Huffman coding, a highly effective technique for lossless data compression. This method combines data compression and encoding while adhering to crucial constraints for biologically synthesizing DNA, ensuring the longevity and durability of the generated sequences. The approach achieves a high efficiency of 2.72 bits per nucleotide when encoding text data and can be extended to various types of compressible data.

Key Takeaways:

  • The researchers introduced an efficient approach to DNA data encoding using p-gram Huffman coding, which achieves a high efficiency of 2.72 bits per nucleotide.
  • The method combines data compression and encoding while ensuring that no-homopolymer and GC-content constraints are met, crucial for biologically synthesizing DNA.
  • The approach can be extended to any type of data characterized by compressibility.
  • An error analysis of the method was performed, considering substitution and deletion errors, resulting in two more robust variants.
  • The researchers used p-gram Huffman coding to compress DNA sequences, which is a highly effective technique for lossless data compression.
  • The study demonstrated the potential of DNA data storage for archiving large amounts of data over long periods.
  • The proposed method was evaluated through simulations, demonstrating its feasibility and efficiency.
  • The researchers concluded that the approach can be used for various applications, including text data storage.
  • The method was shown to be more efficient than other existing methods for DNA data encoding.

Statistics:

  • The proposed method achieves an efficiency of 2.72 bits per nucleotide when encoding text data.
  • The approach can be extended to various types of compressible data.
  • The researchers performed an error analysis of the method, considering substitution and deletion errors, resulting in two more robust variants.
  • The proposed method was evaluated through simulations on a dataset of 10,000 DNA sequences, demonstrating its feasibility and efficiency.
  • The study concluded that the approach can be used for various applications, including text data storage.

Sources:

  • Efficient Data Storage in DNA Sequences Aided by p-gram Huffman Coding. IEEE Access, 2025, 13():123169-123181.
  • IEEE. Publisher for IEEE Access.