X2J on the firing line

· X2J

I've been getting more than one comment about the X2J project I've introduced here and other projects, specifically XStream and XMLBeans. I wanted to clarify the differences (at least the ones I see) between the frameworks to what I intend X2J to be.

The X2J framework, according to myself and the SourceForge title it holds, will "...performs serialization and validation of POJOs to XML, XML to POJOs, POJOs to XSDs, and XSDs to POJOs using Annotations", i.e. will replace the usage of XML Schema Documents by using Annotations. Let's be clear about it - Using X2J means no XSD files needed - You write your own classes, annotate them, and then you writing and reading your data into and from XML documents. It is a binding tool that doesn't use external XSD files for the binding.

You can read about binding tools on a coverpages article about JAXB. To be brief, they state that a binding tool helps "...(1) unmarshal XML content into a Java representation; (2) access, update and validate the Java representation against schema constraint; (3) marshal the Java representation of the XML content into XML content. JAXB gives Java developers an efficient and standard way of mapping between XML and Java code...".

(Yes, I am aware I am shooting really high with this :-) )

Now, about the other frameworks. This is what I understand these other frameworks are meant to do. XStream is a serialization framework. As they say themselves in their FAQ:


Can XStream generate classes from XSD?


No. For this kind of work a data binding tool such as XMLBeans is appropriate.


XMLBeans however is a data binding tool. The way it achieves it is similar to Java Architecture for XML Binding - JAXB. It uses an XSD compiler - scomp - to generate a .jar file containing their pre-compiled XML Schema Object model. Obviously they do a lot more and in some cases a lot better than JAXB, which can be fully read on their Overview page.

The major difference between XMLBeans and X2J is in the binding method. XMLBeans make use of the XSD files; Whenever these files change, they need to recompile their Schema Object model. While this simple compilation process doesn't frighten most of us who know how to use ant, my problem with it is the writing of the XSD files themselves. They are complicated, redundant (as you have the same data structure in your code already in most cases), and hard to maintain (single file for all the data model, instead of the separate Java classes you already have).

The reason behind X2J was to eliminate the need for XSD files, for their complexity and redudancy, and maybe come up with a better way to do XML binding. I would be happy to contribute and join a larger framework - But some are just working on different paths than what I intend X2J to be.

Hope this clarifies some things. I encourage everyone to stand against my words! (or maybe even suggest cooperation?)

Comments (6)

Glenn
Avah: 1. There is a design decision here I question: annotation. Why bother - its unnecessary - use reflection and observe the transient keyword (i.e. if transient, don't persist). Only if you need something that isn't covered by existing Java functionality, in my humble opinion, should you use annotation. In other words, your POJO can be even plainer - they don't need to be annotated, and everything gets persisted by default. Yes, this won't work for classes that don't implement Serializeable (the big one is Thread and other classes that have native state BUT don't store enough information so that an equivalent class instance can be regenerated - a nasty bug, in my view). Working around this is hard but I submit that you had this problem already. 2. What about having an abstract layer for persistence? In other words, XML is simply one of several possible implementations, with the other, of course, being an RDBMS, but with a caveat: Java is driving, so it will maintain the data model. 3. One problem you will run into is versioning. Handling that robustly is important. I applaud your ambition here, and your recognition that there has to be a better way.
Avah
The reason for annotations is to provide the XSD functionality. For example, "Mandatory" for elements and attribute. Another validation could be sizes of arrays: minSize and maxSize should be noted somehow. Another is having a different name for the element than the getter method's name. Another is instanciating a class instance instead of the interface returned in the getter method. There are many reasons for the annotations. I invite you to view the example code provided on the X2J download site, and to read the many posts I've written on X2J - Just go to the X2J archives. Thanks for the feedback though, I need every word I can get, good and bad!
Glenn
1. I didn't make myself very clear, again, I'm afraid. "mandatory" = the absence of the "transient" keyword (by default, its mandatory - persist it all). Array size is available in the reflection api. I would humbly suggest not supporting renaming: use what Java uses - if a user wants to change the name, they would have to go back to Java and change it there. As far as instantiation goes, use Java's rules - instantiate whatever Java would instantiate. Java has figured all of this out (altho notice I'm not saying it would be easy :) ), so there is standard behaviour out there to reproduce. People would be confused by differences/perceive it as non standard. Really, its all about letting Java drive, and using what the POJO has in terms of its characteristics, as exposed by the Reflection api. Only if reflection doesn't reveal something (and I think it reveals pretty much everything you need, based on a project a colleague started) use annotations. 2. What do you think about adding an abstraction layer so that you can support alternative persistence strategies? In addition to RDBMS, serialized files, text files (there are databases that use text files, for example). There is no such thing as "bad" comments, only different views. :)
Avah
The point of X2J is to let the user use his own classes for the XML work. Users, unfortunately, use interfaces. :) I don't want a user to be constrained to using only concrete classes. I think it's poor design to return a LinkedList and not a Collection, for example. On another case, you must agree that reflection doesn't reveal the size of a LinkedList. And it certainly doesn't reveal your Wanted size, in order to perform validation - Both for Collections And for Arrays. There are other cases where Annotations are needed. Did you read the rest of the posts?
Avah
Oh.. And "mandatory" is different than "transient". Transient tells the serialization not to serialize a field. It does not, however, tells an XML validator whether that field was cruicial for the XML in order to be valid.
Avah
Glenn, I just want to make sure that I didn't sound like I'm directing you to RTFM. I just think that the other posts I wrote show almost every annotation I have written in a more thorough way than I could have done in these comments. If you wish, I'd be more than happy to discuss it to a greater extent on a private chat. You can contact me for that at avah [at] crazyredpanda [dot] com. And thanks again for taking interest!