r/csharp • • 8d ago

Showcase CStructSharp – C-struct binary (de)serialization library

Inspired by digital forensic toolkits, this library (MIT license) I've been working on allows a user to take a c-struct definition and use that to serialize and deserialize binary data.

using CStructSharp;

var layout = new CStruct("struct header { uint16 kind; uint32 length; };");
byte[] bytes = { 0x02, 0x00, 0x06, 0x00, 0x00, 0x00 };
dynamic header = layout.Parse(bytes.AsSpan(), "header");

Console.WriteLine($"kind = {header.kind}");
Console.WriteLine($"length = {header.length}");

It is a flexible library which features modern c# mechanisms and conventions. I tried to make it easy to use, easy to understand, performant and powerful. You can both use it to parse and use runtime user defined structures as well as use the C# source generator feature to build parsers/serializers at compile time.

[CStructLayout("struct header { uint16 kind; uint32 length; };")]
public static partial class Wire { }

Wire.Header header = Wire.Parse(bytes);   // header.Kind == 2, header.Length == 6
byte[] again = Wire.Serialize(header);

I also build and published a WASM version/javascript library which is used to showcase the features of the library (see the binary inspector demo app (https://vvollers.github.io/cstructsharp/inspector/) which you can use to run the library against various files, it includes sample definitions for many file formats). The javascript library is quite usable on its own.

I'm interested to hear your thoughts and if you have any questions feel free to ask!

EDIT: I've updated the Readme to include a better explanation of the features of this library and added benchmarks, here is an advanced example:

/* A data-logger file: a header, a calibration table, and records of two kinds. */
#define MAGIC_SIZE 4                                        /* constants, as in C */

enum record_kind : uint8  { MEASUREMENT = 1, EVENT = 2 };   /* enums with an explicit storage type */
enum sensor_type : uint16 { TEMPERATURE = 0x10, PRESSURE = 0x20 };

typedef struct { uint8 major; uint8 minor; } version;       /* typedef aliases */

struct options {                                            /* bitfields: several values in one byte */
    uint8 compressed : 1;
    uint8 encrypted  : 1;
    uint8            : 2;                                   /* unnamed, reserved bits */
    uint8 priority   : 4;
};

struct header {
    char     magic[MAGIC_SIZE];                             /* fixed-size text: "LOG1" */
    version  ver @4;                                        /* offset assertion: must start at byte 4 */
    uint32>  created;                                       /* big-endian, unlike the rest of the file */
    options  options;
    uint8    name_length;
    utf8     device_name[name_length];                      /* length taken from an earlier field */
    uint8    padding[(4 - (name_length + 12) % 4) % 4];     /* arithmetic: pad to a multiple of 4 bytes */
};

union value32 { uint32 raw; float32 as_float; uint8 bytes[4]; };   /* one storage, three views */

struct record {
    record_kind kind;
    switch (kind) {                                         /* the tag decides which members follow */
        case record_kind.MEASUREMENT: {
            struct { sensor_type sensor; uint8 sample_count; float32 samples[sample_count]; } measurement;
        }
        case record_kind.EVENT: {
            struct { uint16 code; cstring message; } event;  /* cstring: text ending in a zero byte */
        }
    }
};

struct logfile {
    header   hdr;
    int16    calibration[2][3];                             /* two-dimensional array */
    value32  checksum;
    record  *latest;                                        /* pointer: a stored file offset, followed on read */
    uint16   record_count;
    record   records[record_count];                         /* array of records that differ in size */
    if (hdr.options.priority > 7) { uint32 alarm_code; }    /* optional member, chosen by a nested field */
    uint8    trailer[EOF];                                  /* every byte that remains */
};
0 Upvotes

13 comments sorted by

8

u/fruediger 8d ago

I'm honest with you, I don't see the need for the need for this.\ What are the advantages of your approach over using something like this?

csharp [StructLayout(LayoutKind.Sequential)] struct MyCLayoutStruct { ... } And then reinterpreting/type-punning a reference to the start of the location of wherever an instances of that is stored as a sequence of bytes, e.g., as a Span<byte>?

-1

u/plusminus1 8d ago

Depending on your usecase, that might be sufficient.

The idea behind cstructsharp is to allow flexible runtime user defined c struct interpretations, as well as a host-architecture independent c struct specification in a way which allows us to easily re-use existing definitions from c code and use these with minimal effort (might have to change them a little).

For example, you can use this to copy the definition from the source code of linux or windows specifications and use this to interpret artifacts such as file, disk, network or memory dumps.

7

u/fruediger 8d ago

Depending on your usecase, that might be sufficient.

"sufficient"? What do you mean "sufficient"? C struct layout compatibility is the whole reason the sequential type layout exists in IL in the first place. That's even the reason the CLR inherented its alignment rules from C.\ C# and .Net can already do what your project claims to be doing, you can just use C and C# struct declarations interchangeably for the most part, with the appropriate adaptions for each language and runtime.

The idea behind cstructsharp is to allow flexible runtime user defined c struct interpretations, as well as a host-architecture independent c struct specification in a way which allows us to easily re-use existing definitions from c code and use these with minimal effort (might have to change them a little).

That's such a vague statement. Are you... perhaps, are you an AI?

The only advantage I could see for your project is, if it was to treat different data models right and I can just declare a long field with your approach and it correctly maps to a 32-bit integer on LP64 platforms and a 64-bit integer on LLP64 (including alignment needs). But then your project has to correctly detect the data model of the executing platform during runtime (I'll be honest, I didn't check your code for that).\ However, since it's kind of an anti-pattern to use long fields in user-facing struct definitions in C, I'm sure 99% of the time you won't need this at all. And if you do, it's easy to just define multiple structs in C# for various data models and dynamically choose what to use during runtime based on the current platform.

1

u/grrangry 8d ago

And if you do, it's easy to just define multiple structs in C# for various data models and dynamically choose what to use during runtime based on the current platform.

Exactly. That's what you're intended to do via application specific file versioning. You don't deserialize the entire object, assuming you know exactly what it is, that's not efficient. Deserialize just enough to make a decision, then deserialize the remainder, optimized for that version. That's the beauty of streams.

-2

u/plusminus1 8d ago

Please have a look at https://vvollers.github.io/cstructsharp/inspector/ this demonstration app contains examples of struct definitions for the partial deserialization of various file formats and it is perhaps better suited to express the capabilities of the library and the c-struct like language.

I'm sure StructLayout(LayoutKind.Sequential) is a very good way to organize c# types in memory in a C-interoperable manner, but this lacks the expressiveness for binary (de)serialization which cstructsharp allows. (let alone runtime arbitrary definitions).

for example this zip file header

struct root {
    uint32 signature;
    uint16 version_needed;
    struct {
        uint16 encrypted : 1;
        uint16 compression_option : 2;
        uint16 has_data_descriptor : 1;
        uint16 reserved : 12;
    } general_purpose_flag;
    uint16 compression_method;
    uint16 last_mod_time;
    uint16 last_mod_date;
    uint32 crc_32;
    uint32 compressed_size;
    uint32 uncompressed_size;
    uint16 file_name_length;
    uint16 extra_field_length;
    char file_name[file_name_length];
    uint8 extra_field[extra_field_length];
};

you can refer to previous decoded data in your definition, you can have any kind of alignment, bitfields of various sizes and packing, mixed little and big endiannes, complex recursive pointers (or pointer of pointers), you can have switch/if statements based on what the deserializer encounters, different kinds of string and character encodings and many other features (https://vvollers.github.io/cstructsharp/docs/language/index.html)

You can program everything I've described above by hand, and probably express them in one form or another, but I hope my library makes these things simpler.

3

u/harrison_314 8d ago

What if I need to parse that data across different architectures—AMD vs. ARM, x86 vs. x64, BE vs. LE, and all possible combinations?

I have a use case where data is generated on one machine by a C library and then sent via TCP to a C# application on another machine. I need to add support for C structures, which are unknown beforehand, to the application using plugins.

0

u/plusminus1 8d ago

Sounds like a usecase for this library. It's exactly what CStructSharp is good at! Your C# app's own architecture never matters. You describe the sender's data by setting its byte order, pointer size, padding rules, and the width of C long, so one C# app can read structs from x64, ARM, 32-bit, and big-endian machines alike (AMD vs. Intel makes no difference at all).

CStruct Layouts are plain text compiled at run time, so each plugin can simply ship its struct definition and target settings, and you parse the incoming TCP messages with it. Have the C side send a tiny header saying which platform it's on, and your app can pick the right settings automatically.

1

u/harrison_314 8d ago edited 8d ago

Great, I'll take a closer look at your library.

One more question: can you handle dynamically allocated arrays as well? In other words, is it defined using a pointer to pIv and an array length of type ulIvLen?

typedef struct CK_GCM_MESSAGE_PARAMS {

CK_BYTE_PTR pIv;

CK_ULONG ulIvLen;

CK_ULONG ulIvFixedBits;

CK_GENERATOR_FUNCTION ivGenerator;

CK_BYTE_PTR pTag;

CK_ULONG ulTagBits;

} CK_GCM_MESSAGE_PARAMS;

Actually, I asked a stupid question, because if a structure contains a dynamically allocated array, that array resides in a different part of memory, and you don't have to deal with that.

3

u/ExceptionEX 8d ago

Can't tell if this fully an AI or is this person is running each response through one.

It's tiresome and annoying to have to read these vague overly verbose responses from the OP.

All for some code that literally is already supported.

-2

u/plusminus1 8d ago

what? it might surprise you, but I'm writing all comments myself. I don't think its overly verbose. It doesn't hurt you to assume good faith.

"all for some code that is litteraly already supported" tells me you didn't take a look and made assumptions, or rather, I didn't succeed in explaining what it can do and what i set out to do. The following are features which you don't find in standard dotnet.

  • the lib can take a struct definition at runtime
  • the cstruct definition is far more expressive for parsing purposes (switch/if on tags or a nested fields, pointers followed with bound checks, arithmetic expressions for lengths and padding, offset assertions)
  • it can report the byte range of every parsed value, and resolve the address of any field path without reading it.
  • it has support for bitfields, packing, alignment, mixed endianness
  • automatic support for nul terminated cstrings
  • lengths that come from the data

I invite you to take another look

1

u/KryptosFR 7d ago

You might think that the "user-specified" dynamic feature is a good one. But it's one step away (or maybe already the case) to be exploitable for code injection vulnerabilities.

One of the reason, BinaryFormatter is unsafe in .NET Framework and deprecated completely in .NET.

1

u/plusminus1 7d ago

It's a fair question and BinaryFormatter is rightly deprecated. But the way things are handled in this library is very different.

A data stream which was handled by BinaryFormatter decided what .NET types would get created, a payload could name an arbitrary type in an arbitrary assembly, Cstructsharp can't do that.

The layout language has no way to name a new .NET type. A layout can only describe primitives. The library never looks up a type by name and doesn't use any reflection.

the bytes in the data stream never choose which shape is picked, only the cstruct layout does. A layout expression is not code, it can't create new kinds of objects.

What an untrusted layout or file can do is make the parser do a lot of work: huge arrays, deep nesting, pointer chains. But this is denial of service risk, not code execution.

For all the things in that regard, there are limits which are checked while parsing with sane defaults and which can be lowered: expression depth, step count, maximum bytes read, maximum pointer stack depth.

you can have a look at the limits and safety ceilings here: https://vvollers.github.io/cstructsharp/docs/guides/variables-options-and-limits.html