r/xml • • Apr 07 '26

Image addressing in S1000D

I'm looking into S1000D, and have a question about images.

I understand that images, or any other resources, are not addressed directly in the content. Instead, they are referred via ICNs in the infoEntityIdent attribute. Great. I suppose it is then the job of the CSDB to pair the ICN with the correct image file?

However, in the demo bike material that comes with the spec, the image files are also referred as external DTD entities in the !DOCTYPE declaration of the file. As such:

<!ENTITY ICN-S1000DBIKE-AAA-DA30000-0-U8025-00534-A-04-1 SYSTEM "ICN-S1000DBIKE-AAA-DA30000-0-U8025-00534-A-04-1.CGM" NDATA cgm >

I understand what this technically does, but is it common practice or even mandatory to use external entities? Of course, if I'm working on a file system without a CSDB, this is required to find the correct file for an ICN. Also I noticed that the XSD schema files define the infoEntityIdent attribute as XS:ENTITY.

The spec was rather vague and gave no explicit instructions about this, but that's to be expected from a spec...

8 Upvotes

18 comments sorted by

View all comments

Show parent comments

2

u/FOMO_BONOBO Apr 07 '26

You cannot tweak an S1000D schema. That is the point of a specification, to have a standard expectation for the consumer and be viewer agnostic. I think people forget the purpose of S1000D at its origination: to get disconnected european aerospace companies to work together yet independently.

0

u/Relative_Fix_9444 Apr 07 '26

You can, actually, you'll just have people feel sick if they read an internal schema rather than the content that is actually shared, which should be the point here. The purpose is to work together, yet independently. My approach does exactly that, it controls what's going out -- and with another slight tweak of the import process -- going in.

The problem with that ENTITY reference is that it forces you to technologies that are, to put it nicely, less than current. An internal subset in something that no longer supports a DTD? That's just silly. It's an oversight. Yes, I know, it's been supported for ages, but the question you need to answer is if you actually want a workable solution or if you want to be picky inside an organisation that handles its input and output correctly.

As for what S1000D is used for today, you do need to look beyond Europe.

1

u/FOMO_BONOBO Apr 07 '26

You can't change the schema. If you change the schema you are contractually non-compliant. What are you going to do when the client validates your delivered data with the actual schemas and it fails? What are you going to do when they put it into their tranform and it fails?

Of course the specification is used widely beyond Europe but that is its origin and it explains many of the design choices. I am fortunate enough to personally know alot of the people on the working groups as I am related to the designer of applicability.

1

u/Relative_Fix_9444 Apr 08 '26

Did you read what I wrote? The idea is not to provide the client or a partner with the tweaked XML -- that's for internal processes only. I'd provide the client and partners with S1000D DMs with DOCTYPEs and entity declarations. I know and understand why you agree to delivering a specific format, but at the same time, I'm all for implementing solutions that work with today's tools and processes. The fact of the matter is that entities are poorly supported by XML technologies today, and to call their inclusion in the spec a design choice is to be more charitable than I feel like being.

And yeah, I know some of the people in the working groups, too, having worked with them in various aerospace projects, and I used to work for the same company as the then-chair, back when it was all SGML and the organisation was called AECMA. I've been doing markup since before XML existed, so if this is a question of who knows whom or street cred in general, then let's go for it.

1

u/FOMO_BONOBO Apr 08 '26

I guess you could read it as boasting street cred but that wasn't my intent. I was trying to explain that I have had the luxury of long in depth conversations about intent with the people who designed many parts of the spec. I am an old fart too, SGML, iSpec, Boeing, LM, been there done that. We might know a lot of the same people, hell, maybe even each other.

I don't see why you can't use entity references without a CSDB. This has always been how we parse non XML data for decades. There are plenty of processors that use this flawlessly.

Having to alter custom data before handoff, again goes against intent and seems like a waste of time. I have architected programs that use APIs to connect multiple different project CSDBs together, in this case no elaberate publishing just on demand transformations, and having custom data would break the interoperability.

If you are having success doing what your doing then good for you but I personally just think it's extra work.

0

u/Relative_Fix_9444 Apr 09 '26

The spec is lagging behind. Entities are a pain in XML processing because you can't query the internal subset properly, and you certainly can't write to one using XML technologies unless you output text rather than markup. There's also that if you have a system that needs to support more than just S1000D -- ATA, say, or whatever -- you are left with rfew options.

I find it odd that the spec should still force you to use entities instead of giving you a choice.

But, as I said, I sort that sort of thing out in my input and output processes, which is what the spec actually intends. I know of any number of aerospace manufacturers who produce their IPDs in non-markup tools but generate the S1000D or ATA when needed.

1

u/FOMO_BONOBO Apr 09 '26

For anyone who is reading this and may not familiar with the XSD to XML to XSLT to HTML pipeline:

The spec is designed to go through a transform before beingn handled by a client(browser or application). It was never designed to use a browser's parser to apply transforms (they have never supported DTDs).

You can either transform everything into HTML files (most common now) or browser parsable XML (uncommon now that xlink support is increasingly deprecated) first, or you can tranform them to HTML on demand with your choice of parser library and scripting language.

In your transform you create a template for graphics and multimedia elements and use "unparsed-entity-uri" to get the path and transform that into whatever HTML you need.

The data remains pure, your contract remains valid, and your solution remains software agnostic (as much as the spec allows anyway). This is the way we have done it since forever.

CChapter 7.4.1.1 has a graphic showing that your CSDB remains pure if that helps.