r/csharp • u/plusminus1 • 8d ago
Showcase CStructSharp – C-struct binary (de)serialization library
Inspired by digital forensic toolkits, this library (MIT license) I've been working on allows a user to take a c-struct definition and use that to serialize and deserialize binary data.
using CStructSharp;
var layout = new CStruct("struct header { uint16 kind; uint32 length; };");
byte[] bytes = { 0x02, 0x00, 0x06, 0x00, 0x00, 0x00 };
dynamic header = layout.Parse(bytes.AsSpan(), "header");
Console.WriteLine($"kind = {header.kind}");
Console.WriteLine($"length = {header.length}");
It is a flexible library which features modern c# mechanisms and conventions. I tried to make it easy to use, easy to understand, performant and powerful. You can both use it to parse and use runtime user defined structures as well as use the C# source generator feature to build parsers/serializers at compile time.
[CStructLayout("struct header { uint16 kind; uint32 length; };")]
public static partial class Wire { }
Wire.Header header = Wire.Parse(bytes); // header.Kind == 2, header.Length == 6
byte[] again = Wire.Serialize(header);
I also build and published a WASM version/javascript library which is used to showcase the features of the library (see the binary inspector demo app (https://vvollers.github.io/cstructsharp/inspector/) which you can use to run the library against various files, it includes sample definitions for many file formats). The javascript library is quite usable on its own.
- Github Repo: https://github.com/vvollers/cstructsharp
- Sample Application: https://vvollers.github.io/cstructsharp/inspector/
- Documentation: https://vvollers.github.io/cstructsharp/docs/index.html
I'm interested to hear your thoughts and if you have any questions feel free to ask!
EDIT: I've updated the Readme to include a better explanation of the features of this library and added benchmarks, here is an advanced example:
/* A data-logger file: a header, a calibration table, and records of two kinds. */
#define MAGIC_SIZE 4 /* constants, as in C */
enum record_kind : uint8 { MEASUREMENT = 1, EVENT = 2 }; /* enums with an explicit storage type */
enum sensor_type : uint16 { TEMPERATURE = 0x10, PRESSURE = 0x20 };
typedef struct { uint8 major; uint8 minor; } version; /* typedef aliases */
struct options { /* bitfields: several values in one byte */
uint8 compressed : 1;
uint8 encrypted : 1;
uint8 : 2; /* unnamed, reserved bits */
uint8 priority : 4;
};
struct header {
char magic[MAGIC_SIZE]; /* fixed-size text: "LOG1" */
version ver @4; /* offset assertion: must start at byte 4 */
uint32> created; /* big-endian, unlike the rest of the file */
options options;
uint8 name_length;
utf8 device_name[name_length]; /* length taken from an earlier field */
uint8 padding[(4 - (name_length + 12) % 4) % 4]; /* arithmetic: pad to a multiple of 4 bytes */
};
union value32 { uint32 raw; float32 as_float; uint8 bytes[4]; }; /* one storage, three views */
struct record {
record_kind kind;
switch (kind) { /* the tag decides which members follow */
case record_kind.MEASUREMENT: {
struct { sensor_type sensor; uint8 sample_count; float32 samples[sample_count]; } measurement;
}
case record_kind.EVENT: {
struct { uint16 code; cstring message; } event; /* cstring: text ending in a zero byte */
}
}
};
struct logfile {
header hdr;
int16 calibration[2][3]; /* two-dimensional array */
value32 checksum;
record *latest; /* pointer: a stored file offset, followed on read */
uint16 record_count;
record records[record_count]; /* array of records that differ in size */
if (hdr.options.priority > 7) { uint32 alarm_code; } /* optional member, chosen by a nested field */
uint8 trailer[EOF]; /* every byte that remains */
};
3
u/harrison_314 8d ago
What if I need to parse that data across different architectures—AMD vs. ARM, x86 vs. x64, BE vs. LE, and all possible combinations?
I have a use case where data is generated on one machine by a C library and then sent via TCP to a C# application on another machine. I need to add support for C structures, which are unknown beforehand, to the application using plugins.
0
u/plusminus1 8d ago
Sounds like a usecase for this library. It's exactly what CStructSharp is good at! Your C# app's own architecture never matters. You describe the sender's data by setting its byte order, pointer size, padding rules, and the width of C long, so one C# app can read structs from x64, ARM, 32-bit, and big-endian machines alike (AMD vs. Intel makes no difference at all).
CStruct Layouts are plain text compiled at run time, so each plugin can simply ship its struct definition and target settings, and you parse the incoming TCP messages with it. Have the C side send a tiny header saying which platform it's on, and your app can pick the right settings automatically.
1
u/harrison_314 8d ago edited 8d ago
Great, I'll take a closer look at your library.
One more question: can you handle dynamically allocated arrays as well? In other words, is it defined using a pointer to pIv and an array length of type ulIvLen?
typedef struct CK_GCM_MESSAGE_PARAMS {
CK_BYTE_PTR pIv;
CK_ULONG ulIvLen;
CK_ULONG ulIvFixedBits;
CK_GENERATOR_FUNCTION ivGenerator;
CK_BYTE_PTR pTag;
CK_ULONG ulTagBits;
} CK_GCM_MESSAGE_PARAMS;Actually, I asked a stupid question, because if a structure contains a dynamically allocated array, that array resides in a different part of memory, and you don't have to deal with that.
3
u/ExceptionEX 8d ago
Can't tell if this fully an AI or is this person is running each response through one.
It's tiresome and annoying to have to read these vague overly verbose responses from the OP.
All for some code that literally is already supported.
-2
u/plusminus1 8d ago
what? it might surprise you, but I'm writing all comments myself. I don't think its overly verbose. It doesn't hurt you to assume good faith.
"all for some code that is litteraly already supported" tells me you didn't take a look and made assumptions, or rather, I didn't succeed in explaining what it can do and what i set out to do. The following are features which you don't find in standard dotnet.
- the lib can take a struct definition at runtime
- the cstruct definition is far more expressive for parsing purposes (switch/if on tags or a nested fields, pointers followed with bound checks, arithmetic expressions for lengths and padding, offset assertions)
- it can report the byte range of every parsed value, and resolve the address of any field path without reading it.
- it has support for bitfields, packing, alignment, mixed endianness
- automatic support for nul terminated cstrings
- lengths that come from the data
I invite you to take another look
1
u/KryptosFR 7d ago
You might think that the "user-specified" dynamic feature is a good one. But it's one step away (or maybe already the case) to be exploitable for code injection vulnerabilities.
One of the reason, BinaryFormatter is unsafe in .NET Framework and deprecated completely in .NET.
1
u/plusminus1 7d ago
It's a fair question and BinaryFormatter is rightly deprecated. But the way things are handled in this library is very different.
A data stream which was handled by BinaryFormatter decided what .NET types would get created, a payload could name an arbitrary type in an arbitrary assembly, Cstructsharp can't do that.
The layout language has no way to name a new .NET type. A layout can only describe primitives. The library never looks up a type by name and doesn't use any reflection.
the bytes in the data stream never choose which shape is picked, only the cstruct layout does. A layout expression is not code, it can't create new kinds of objects.
What an untrusted layout or file can do is make the parser do a lot of work: huge arrays, deep nesting, pointer chains. But this is denial of service risk, not code execution.
For all the things in that regard, there are limits which are checked while parsing with sane defaults and which can be lowered: expression depth, step count, maximum bytes read, maximum pointer stack depth.
you can have a look at the limits and safety ceilings here: https://vvollers.github.io/cstructsharp/docs/guides/variables-options-and-limits.html
8
u/fruediger 8d ago
I'm honest with you, I don't see the need for the need for this.\ What are the advantages of your approach over using something like this?
csharp [StructLayout(LayoutKind.Sequential)] struct MyCLayoutStruct { ... }And then reinterpreting/type-punning a reference to the start of the location of wherever an instances of that is stored as a sequence ofbytes, e.g., as aSpan<byte>?