Added — Claude Code skill
-
A Claude Code skill for odmlib ships in the repository at
.claude/skills/odmlib/.
It teaches Claude the loader-per-standard mapping, namespace registration, element
ordering, the three validation layers, and the serialization pitfalls that are easy to
get wrong by hand. Contents:SKILL.md, fourreferences/*.md(API reference, models,
validation, Dataset-JSON), and eight runnableexamples/*.py.
.claude/skills/odmlib.skillis the same tree packed as a zip for distribution. -
Repo-only — it is not part of the PyPI package.
pip install odmlibdoes not install
the skill;[tool.setuptools.packages.find]includesodmlib*only. Install it by copying
.claude/skills/odmlib/into a project's.claude/skills/, or into~/.claude/skills/
to make it available everywhere. See the Claude Code Skill sections ofREADME.mdand
CLAUDE.md. -
It describes odmlib 0.2.1 and later. Several behaviors it documents do not hold on
0.2.0 — namespace-awareto_xml_string(), opt-in context-manager writing, and full error
enumeration undercollect_errors=True. The skill advises detecting capabilities rather
than comparing version strings, since a pre-release sorts below its release under PEP 440. -
Guarded by tests, not just prose.
tests/test_skill_contract.pypins the skill's
documented signatures, front-matter limits, import surface, schema pairs, and the
list-vs-object shape rule against live introspection;tests/test_skill_examples.pyruns
all eight examples;tests/test_skill_bundle.pychecks the packed.skillbundle matches
the source tree by SHA-256. Repack withpython scripts/build_skill_bundle.pyafter
editing any skill file. These run in a single CI cell. -
Feedback welcome via the
skill-feedback.ymlissue template ("Claude generated
incorrect odmlib code"), which captures the prompt, the generated code, and the versions
involved.
Added — to_element()
-
ODMElement.to_element()returns a standard, namespace-resolved ElementTree Element.
Getting a tree out of odmlib previously meantto_xml(), which builds the library's
internal serialization buffer: prefix-literal tags (def:leaf, not Clark notation) and
noxmlnsdeclarations. That tree cannot be re-parsed on its own for Define-XML
(ParseError: unbound prefix), lands in no namespace for ODM, fails
ET.canonicalize(), and does not support namespace-awarefind().to_element()
returns a tree parsed fromto_xml_string(), so all of those work:elem = define.Study.MetaDataVersion.ItemGroupDef[0].to_element() elem.find("{http://www.cdisc.org/ns/def/v2.1}leaf") # namespace-aware find host = ET.Element("SubmissionPackage"); host.append(elem) # embeds correctly ET.indent(ET.ElementTree(elem)) # pretty-prints and re-parses
It costs one serialize + reparse — about 8 ms for a 166 KB Define-XML document.
DatasetJSONElement.to_element()raisesNotImplementedError, matching its sibling
to_xml/to_xml_string/write_xmloverrides. Pinned by
tests/test_xml_string_serialization.py::TestToElement.to_xml()is unchanged and is not deprecated — it remains the shared tree builder
behindto_xml_string()andwrite_xml(), and existing callers keep working. It has
been demoted in the documentation from the head of the serialization list to a
trailing "internal tree builder" entry; preferto_element()when you want a tree.One caveat: re-serializing a
to_element()tree withET.tostring()picks prefixes
from ElementTree's process-globalregister_namespace()map, so the prefix spelling
may differ from the source (odm:ODMrather than a defaultxmlns). The namespaces
are identical. Useto_xml_string()orwrite_xml()when exact output matters.
Added — to_xml_string(xml_declaration=True)
-
ODMElement.to_xml_string()accepts a keyword-onlyxml_declarationflag.
The string path previously had no way to emit an XML declaration, so callers
handing a string to a consumer that requires one had to prepend it by hand —
and comparing a string against awrite_xml()file silently disagreed on the
first 39 bytes. The two paths now line up exactly:odm.to_xml_string() # bytes write_xml() writes AFTER <?xml ...?> odm.to_xml_string(xml_declaration=True) # exactly what write_xml() writes
The default stays
False— this string is the documented input to
ODMLoader.load_odm_string()and 0.2.0 shipped it declaration-free, so
flipping it would silently change output for every existing caller. The
parameter is keyword-only.dataset_json_1_1.model.DatasetJSON.to_xml_string()
accepts the same keyword so it still raises the intendedNotImplementedError
rather than aTypeError. Pinned by
tests/test_xml_string_serialization.py::TestXmlDeclarationOption.
Fixed — write_xml() emitted CRLF line endings on Windows
write_xml()now writes LF line endings on every platform. It passed a filename to
ElementTree.write(), which opens the file in text mode; on Windows that translated every
\nto\r\n, so the file no longer matchedto_xml_string()byte for byte and the
documented equivalence between the two was false there. The writer now opens a binary
handle itself. XML written on Windows changes from CRLF to LF, which makes checksums and
diffs portable across platforms.
Fixed — Dataset-JSON file I/O used the platform's default encoding
DatasetJSON.write_json(),write_ndjson(),read_json()andread_ndjson()now pin
encoding="utf-8". They relied on the locale default, which is UTF-8 on Linux and macOS
but cp1252 on most Windows installs — so a Dataset-JSON file containing non-ASCII text and
produced by another tool could decode incorrectly or raise. Files odmlib itself writes were
already ASCII-safe (json.dumpsescapes non-ASCII by default), so existing output is
unaffected.
Fixed — nested elements serialized with the wrong namespace
-
A nested element reached by walking a loaded tree now serializes with the
namespaces its document was loaded under. The per-document namespace
snapshot was bound only to the objects the loader returns directly
(root(),Study(),MetaDataVersion(),create_odmlib()). Anything
reached by walking —define.Study.MetaDataVersion— had no snapshot and fell
back to current global registry state, so merely importing a second Define
model package changed its output:import odmlib.define_2_0.model # re-registers def: -> v2.0 globally define.to_xml_string() # root: xmlns:def=".../def/v2.1" (right) define.Study.MetaDataVersion.to_xml_string() # nested: ".../def/v2.0" (WRONG)
ns_registry.bind_document_namespaces()takes a newrecursive=False
parameter, andloader.ODMLoader._bind_namespaces()passesrecursive=True,
which covers all four loader entry points and theopen_odm/open_define
context managers.write_xml()on a nested element is fixed by the same
change. Measured cost: 1.4 ms to bind all 2 086 elements of a 136 KB
Define-XML file, sharing one snapshot dict. Pinned by
tests/test_xml_string_serialization.py::TestNestedElementNamespaceBinding.Residual limitation: an element constructed after the load and grafted
in still carries no snapshot and uses global state. Bind it explicitly:import odmlib.ns_registry as NS NS.bind_document_namespaces(new_elem, NS.get_document_namespaces(root))
Deprecated — NamespaceRegistry.set_odm_namespace_attributes_string()
- Emits
OdmlibDeprecationWarning; will be removed in 0.3.0. Since 0.2.1
to_xml_string()declares its own namespaces, which makes this string-patching
helper a no-op on any string it would normally be given. It has no callers in
odmlib. Remove the call; no replacement is needed.
Changed — a misnamed serialization test
test_schema_ordered_serialization.py::test_to_xml_string_round_trip_unchanged
never calledto_xml_string()— it used rawET.tostring(), which is false
coverage of exactly the path that went untested. Renamed to
test_to_xml_element_order_survives_reparse(its body is a valid element-order
test) and a realto_xml_string()round-trip added beside it as
test_to_xml_string_round_trip_preserves_order.
Added — ARM 1.0 XSD schema validation
-
ARM documents can now be schema-validated against a bundled CDISC XSD.
Previously odmlib shipped no ARM schema, so validating Analysis Results
Metadata meant supplying your own viaxsd_file=. Two schema sets are now
bundled underodmlib/schemas/arm/, registered in
schema_manager._MAIN_SCHEMAand reachable through the existing
ODMSchemaValidatorAPI:from odmlib.odm_parser import ODMSchemaValidator validator = ODMSchemaValidator(standard="arm", version="1.0-define2.1") validator.validate_file("define-adam.xml")
(standard, version)Validates ("arm", "1.0")ARM 1.0 in a Define-XML 2.0 document (CDISC original) ("arm", "1.0-define2.1")ARM 1.0 in a Define-XML 2.1 document Two sets are required because ARM layers onto Define-XML, and Define-XML 2.0
and 2.1 use differentdef:namespace URIs — the pairings are not
interchangeable. Use"1.0-define2.1"withodmlib.arm_1_0, which extends
odmlib.define_2_1. The1.0-define2.1schema set is derived by odmlib from
the CDISC ARM 1.0 schema by retargeting the Define-XML dependency; element and
type declarations are unchanged from the original.Both ARM schemas are supersets of their base Define-XML schema, so either also
validates an ARM-free Define-XML document of the matching version.
Fixed — ARM AnalysisResult serialized in a schema-invalid element order
arm_1_0.model.AnalysisResultdeclared its child elements in the wrong
order, so every ARM document odmlib wrote failed XSD validation. Descriptor
declaration order is serialization order, and the ARM schema requires the
sequenceDescription, AnalysisDatasets, Documentation, ProgrammingCode;
the model declaredAnalysisDatasetslast.AnalysisDatasetshas been moved
ahead ofDocumentation. Reading ARM documents was unaffected — only output
was wrong, which went unnoticed while no ARM XSD was bundled to check it.
Fixed — collect_errors=True now collects every error
-
validate(collect_errors=True)returned at most three errors. It wrapped
each of its three validation layers in a singletry/except, and every layer
was itself fail-fast, so the returned list held at most one error per layer
no matter how broken the document was. A document with 50 misordered elements
reported 1. Each layer now enumerates every problem it finds
(odm_element.py,oid_generator.py,exceptions.py):- Order: one error per misordered element; the walk now recurses into the
children of a misordered element instead of aborting. - OID: one error per duplicate OID and per bad reference. A duplicate no
longer aborts the traversal, so the reference checks — which previously
never ran at all once a duplicate was found — now execute. On a duplicate
the first definition is kept. - Conformance: the bundled Cerberus result is expanded into one
OdmlibConformanceErrorper failing field, each with a dottedfield_path
and the complete raw dict still oncerberus_errors.
Behaviour change:
len(errors)may now be larger than before for the
same document, and errors are no longer at predictable list positions —
filter by exception type rather than by index. Fail-fast mode
(collect_errors=False, the default) is unchanged. - Order: one error per misordered element; the walk now recurses into the
Added
validate(max_errors=N): caps collection on a badly broken document.
Validation stops the moment the cap is reached — enforced inside each layer,
not by truncating afterwards — and a finalOdmlibErrorLimitErroris
appended, so the list holds at mostN + 1entries. Defaults toNone
(uncapped).- Collecting-checker protocol (
odmlib.exceptions): theErrorReporting
mixin (report()/collecting()) and theis_collecting_checker()
capability check.DynamicOIDRefimplements it; the deprecated manual
OIDRefclasses and duck-typed custom checkers do not and degrade gracefully
to one error for the OID layer.verify_oids()also collects when a sink is
installed viawith checker.collecting(collector):. DynamicOIDRef.reset(): clears accumulated OID state so one checker can
validate a second document.validate()now warns in collect mode when
handed a checker that still holds state from a previous run.flatten_cerberus_errors()andOdmlibConformanceError.expand()/
.field_pathfor per-field conformance reporting.ErrorCollectoracceptsmax_errorsand exposesis_full/truncated.
Uncapped collectors never raise, so existing usage is unaffected.
Changed
DynamicOIDRef.check_oid_refsiterates its reference sets in sorted order.
Previously which bad reference was reported first varied with
PYTHONHASHSEED; error ordering is now deterministic in both modes.
Fixed — validation correctness (code-review remediation)
- Cerberus schema isolation: each
MetadataSchemanow uses a private
SchemaRegistryinstead of the process-globalcerberus.schema_registry.
Previously, instantiating checkers for two model versions (e.g. ODM 1.3.2
and Define-XML 2.1) silently corrupted each other's schemas — the last
checker instantiated won for every shared schema name, rejecting valid
documents and accepting invalid ones. (*/rules/metadata_schema.py) - Define-XML leaf references are now validated:
leaf/@IDdefinitions and
DocumentRef/@leafID/ItemGroupDef/@def:ArchiveLocationIDreferences are
surfaced to the OID checkers. A danglingleafIDpreviously passed
verify_oids()silently. (odm_element.py,oid_generator.py) - Duplicate OIDs on skip-listed elements are now detected:
DynamicOIDRef.add_oidchecks uniqueness before honouringskip_elem, so
duplicateItemGroupDefOIDs in Define-XML are caught (skip_elem now only
exempts an element from reference-target checking). - Deprecated manual
OIDRefcrash guards:add_oid_refno longer raises
KeyErroron unregistered attributes (e.g.SignatureOID), and
check_unreferenced_oidsno longer raisesKeyErrorfor element types
missing fromdef_ref(all three model packages). - Conformance schema drift:
Presentationadded to the ODM 1.3.2
MetaDataVersioncerberus schema (valid documents were rejected as
"unknown field");Repeatingis nowrequiredfor
StudyEventDef/FormDef/ItemGroupDef, matching the model and the spec. - Valueset lookup follows inheritance:
ValidValuesresolves the
ClassName.attrvalueset key via the MRO, so subclasses of model classes
keep their parents' valueset validation.
Fixed — namespaces and serialization
- Per-document namespaces: documents loaded via
ODMLoaderremember the
namespace registry state they were loaded under;write_xml()and
to_xml_string()use that snapshot, so loading a second document (e.g.
ODM 2.0 after ODM 1.3.2) no longer changes thexmlnsa previously loaded
document serializes with. (ns_registry.py,loader.py,odm_element.py) to_xml_string()output is namespace-well-formed: it now includes
xmlnsdeclarations (previously prefixed tags likedef:ValueListDef
had no declaration anywhere, so the string could not be re-parsed).
set_odm_namespace_attributes_string()is a no-op on such strings.- Only used prefixes are declared: serialization emits
xmlns:entries
only for prefixes actually present in the tree, so importing an unrelated
model package (arm/ct/dataset) no longer pollutes output; the redundant
xmlns:xmldeclaration is gone (thexmlprefix is reserved). - ARM namespace registration:
arm_1_0no longer registers itself as the
default namespace (import-order dependent corruption); the registry now
keeps a single default (a new default replaces the previous one), and an
empty registry raisesOdmlibNamespaceErrorinstead ofIndexError.
Fixed — parsing and loading
- Security — DOCTYPE rejection: XML parsing rejects documents containing a
DOCTYPE declaration (billion-laughs / entity-expansion DoS defense) via a
cheap expat prolog pre-scan; ODM never requires DTDs. (odm_parser.py) - Clear parse errors: malformed XML/JSON and unknown root elements now
raiseOdmlibParsingError(with hints) instead of rawParseError,
JSONDecodeError, orAttributeError(all six loaders + parser). - Encoding: JSON reads use
utf-8-sig(BOM tolerant) and JSON/XML writes
use UTF-8 explicitly, instead of the platform default encoding. - ODM 2.0 clinical data parsing:
ODMParser.AdminData/ClinicalData/ ReferenceDatahonour the configured namespace registry instead of a
hardcoded ODM v1.3 URI (they silently returned[]for ODM 2.0 documents).
Fixed — object model
- Auto-created children no longer leak into output: reading an unset
optional child element still auto-creates it (the
rc.ErrorMessage.TranslatedText.append(...)idiom is preserved) but
serialization skips auto-created elements that were never populated —
read-only inspection previously injected spurious empty elements (e.g.
<BasicDefinitions/>) into XML/JSON output. Reading a child whose class
has required attributes returnsNoneinstead of relying on the
deprecatedValueErrorbase ofOdmlibRequiredAttributeError. find/find_all/find_byguards: searching an unset single child or
a scalar attribute name returnsNone/[]instead of raising
AttributeError.- Restricted subclasses are now consistent: assigning a field that a
subclass deliberately dropped (e.g.Questionon a Define-XMLItemDef)
raisesOdmlibTypeErrorinstead of silently storing a value that
serialized inconsistently or not at all; constructor kwarg checking and
the required-attribute check now use the class's effective field set. - Single-child type validation: an
ODMObjectdescriptor validates the
items when a list is assigned (previously ANY list was accepted unchecked). - ODMBuilder scope pointers:
add_study/add_metadata_version/
add_item_group_defclear stale current-element pointers, so
with_description()/with_alias()no longer attach to a previous
ItemDef/MetaDataVersion after a new scope opens. - Converter dataset-name collisions:
dataset_xml_to_dataset_jsonwarns
and keys a colliding dataset by its fullItemGroupOIDinstead of
silently overwriting (e.g.IG.AEvsSUPP.AEboth deriving "AE"). - DataFrame row drops are visible:
dataframe_to_itemswarns (with row
index and reason) for rows that fail element construction instead of
silently returning fewer elements.
Added
merge_fields=Trueclass keyword for model subclassing: a subclass
declared asclass MyItemDef(ODM.ItemDef, merge_fields=True)inherits all
base-class fields/elements without redeclaring them (redeclaring a field
moves it to the subclass position). The default remains the historical
declare-from-scratch behaviour that the Define-XML models use to restrict
inherited ODM fields. (ODMMeta)- Context managers are read-only by default (breaking):
open_odm()/
open_define()without anoutput_fileno longer rewrite the input file
on exit. Writing requires an explicitoutput_file, orwrite_on_exit=True
to opt in to an in-place update.
Performance
- Format-validation regexes (datetime/partial/incomplete/SAS names) are
compiled once at import instead of on every attribute assignment. - Compiled XML schemas are cached by path (
ODMSchemaValidatorno longer
recompiles the XSD per instance); cerberusValidatorobjects are cached
perMetadataSchemainstance. - Removed no-op filter-dict allocations from the
to_dict/OID/order
traversals; Define manual checkers setis_verifiedso
unreferenced_oids()no longer re-runs the full verification walk;
dataframe.pyavoidsiterrows.
Fixed — ODM v2.0 model/XSD alignment (the five structural gaps)
v0.2.0 shipped with five ODM 2.0 structural features that produced schema-invalid output
if used, listed under Known Limitations in that release as deferred to v0.2.1. All five
are closed. tests/test_odm_2_0_known_gaps.py kept its assertions as regression guards
after the xfail(strict=True) markers came off, so none of them can quietly reopen.
ConditionDefgained the XSD-requiredMethodSignature, and itsDescription
became required. EveryConditionDefodmlib built was previously schema-invalid, and so
transitively was anyCollectionExceptionConditionOIDpointing at one.FormalExpressionbecame element-based. See the breaking change below.Protocolno longer carriesStudyEventRef. The ODM 2.0 XSD reaches study events
throughStudyEventGroupRef→StudyEventGroupDef, soProtocolgained
StudyEventGroupRef,StudyTimingsandWorkflowRefinstead.MetaDataVersion.StudyTimingmoved toProtocol/StudyTimings, its XSD position. A
newStudyTimingscontainer holds theStudyTimingelements, which makes the four
timing-constraint classes reachable in a valid document for the first time.StudyEventGroupDefgained its required child group —StudyEventGroupRefand
StudyEventRef, plus the optionalWorkflowRefandCoding. It previously could not
satisfy its own content model at all.
Breaking within draft ODM 2.0: FormalExpression no longer takes _content. The XSD
models the expression as a choice of exactly one Code (inline source) or
ExternalCodeLib (a reference to an external library), not as element text. Code that
constructed FormalExpression(Context=..., _content="...") under model_package="odm_2_0"
now raises OdmlibTypeError; the expression text moves into a Code child:
# before (schema-invalid)
FormalExpression(Context="Python", _content="age >= 18")
# after
FormalExpression(Context="Python", Code=Code(_content="age >= 18"))ODM 1.3.2 and Define-XML FormalExpression are unchanged and remain text-based.
ODMBuilder.add_method_def(formal_expression=...) and add_condition_def(...) still take
plain text and wrap it correctly for whichever model package is in use, so builder callers
need no change.
Two follow-on gaps found while verifying the five also closed: CommentDef and Leaf
became MetaDataVersion children — both classes already existed but were unreachable from
a document, so a CommentOID could never resolve — and DocumentRef's leafID attribute
was corrected to LeafID, the spelling the XSD requires.
Changed — ODM v2.0 model/XSD alignment (phase 1)
Comparing odmlib/odm_2_0/model.py against the bundled ODM 2.0 XSD mechanically —
rather than by hand, one document at a time — surfaced far more divergence than the
structural gaps closed earlier in this release. tests/test_odm_2_0_xsd_alignment.py
now pins the whole divergence set against an allowlist, so drift cannot be introduced
or silently fixed without the test failing. ODM_XSD_ALIGNMENT.md plans the rest.
This first pass corrects the members that made odmlib's own output schema-invalid.
These are breaking changes within draft ODM 2.0. There is no alias layer in odmlib —
a descriptor name is the XML attribute name — so a document using an old spelling now
raises OdmlibTypeError on load in strict mode, and loses the value silently in
permissive mode. That is the intended outcome: none of the old spellings were ever valid
against the ODM 2.0 schema. ODM 1.3.2, Define-XML, ARM, CT and Dataset-XML/JSON are
untouched.
Telecom.value→Telecom.Value.TelecomAttributeDefinitioncapitalises it; the
lower-case spelling made every document containing aTelecomschema-invalid. Same
class of bug as theDocumentRefleafID→LeafIDfix.ODM.Archivalremoved. It is not inODMAttributeDefinition. TheODM.Archival
value-set key went with it.User.Prefix/User.Suffixare now child elements, matching the XSD, and new
Prefix/Suffixclasses back them.User.DisplayNameremoved — not part of the
ODM 2.0 XSD — along with theDisplayNameclass.Organizationreplaced. The text-only leaf carried over from ODM 1.3.2 was
orphaned — referenced by nothing — and has been replaced by the element the XSD
defines:OID/Name/Typerequired, plusRole,LocationOID,
PartOfOrganizationOIDandDescription/Address/Telecomchildren. It is now an
AdminDatachild in its own right; aUserlinks to one throughOrganizationOID,
which consequently resolves underverify_oids()for the first time.RelativeTimingConstraint: the fourPredecessor*/Successor*OID attributes were
replaced by the XSD'sPredecessorOIDandSuccessorOID. Either may reference any
structural element, so splitting them by event-vs-group was both wrong and unnecessary.
oid_generator_config.pyskips the two new names in their place.TransitionTimingConstraint:TimepointRelativeTarget→TimepointTarget, and the
missingTypeattribute was added.Originchildren reordered to the XSD sequence —Description,SourceItems,
DocumentRef. odmlib serializes in declaration order, so the old order emitted
DocumentReffirst and the document failed validation. This changes serialized output
for anyOriginthat carries aDocumentRef.WorkflowEndgained_content. Its XSD type isxs:simpleContentovertextand
the model had no text member at all.DurationDateTimeStringaccepts the wholedurationDatetimeunion. It enforced
[+-]P{n}Wonly, rejecting XSD-valid values such asP3D,PT1H30Mand the empty tag.
The descriptor is used solely by the four ODM 2.0 timing-constraint classes, so no other
standard is affected.DocumentRef/@LeafIDis ref-checked again. ID-based reference detection is
hardcoded string sniffing, and neither the attribute name nor the lower-caseleaf
target ever matchedodm_2_0.LeafIDnow resolves toLeaf, so a dangling document
reference raises like any other unresolved OID.- New value-set keys for
Organization.TypeandTransitionTimingConstraint.Type.
Changed — ODM v2.0 model/XSD alignment (phase 2)
Required flags and cardinalities brought into line with the ODM 2.0 XSD. Most of this
is declarative — required on a child-element descriptor is not enforced at
construction — but three groups do change behaviour.
Attributes the model demanded that the XSD marks optional no longer raise when
omitted: FormalExpression.Context, MethodDef.Type, TargetTransition.ConditionOID,
and the pre/post window attributes on all four timing constraints
(AbsoluteTimingConstraint, RelativeTimingConstraint, TransitionTimingConstraint,
DurationTimingConstraint). This is a pure loosening; existing code that supplies them
is unaffected.
Standard.Status is now required, matching use="required" in the XSD. Unlike the
element-level tightenings this is enforced at construction, so
Standard(OID=…, Name=…, Type=…, Version=…) without a Status now raises
OdmlibRequiredAttributeError. Leaf.Title, MethodDef.MethodSignature and
Study.MetaDataVersion were also marked required, but as child elements those are
declarative only.
Cardinality corrections change the shape of four attributes. Code that indexed or
appended to them needs updating:
- Now single (
maxOccurs="1"in the XSD):MetaDataVersion.Standards,
Address.StreetName, andWorkflowRefonItemGroupDef,StudyEventDefand
StudyStructure. Note the severalStandardentries live inside the oneStandards
container, so nothing is lost. - Now lists (
maxOccurs="unbounded"):AnnotatedCRF.DocumentRef,
SupplementalDoc.DocumentRefandStudyTiming.TransitionTimingConstraint. Each could
previously hold only one reference where the schema allows many.
Assigning a list to a now-single child is still tolerated by ODMObject.__set__ and
serializes every item, so odmlib does not reject over-long content itself — XSD
validation is what catches it.
Added — ODM v2.0 model/XSD alignment (phase 3)
Members the ODM 2.0 XSD defines but the model never had. All additive — no existing
attribute changed name, type or cardinality.
New classes: Class and SubClass (ODM 2.0 models a dataset's general observation
class as a child element with a Name attribute, not the Define-XML-style
ItemGroupDef/@Class attribute), ValueListRef, HouseNumber and GeoPosition.
New attributes: ItemGroupDef.Structure and .ArchiveLocationID,
ItemGroupRef.MethodOID, CodeListItem.ExtendedValue, Include.href,
PDFPageRef.Title, RangeCheck.ItemOID, Location.OrganizationOID.
New children: MetaDataVersion gains AnnotatedCRF, SupplementalDoc,
ValueListDef and WhereClauseDef — all four classes already existed but were
unreachable from a document, the same gap CommentDef and Leaf had. ItemDef gains
ValueListRef; ItemGroupDef gains Class and Leaf; MethodDef gains DocumentRef;
RangeCheck gains MethodSignature; Origin, SourceItems and StudyEventDef gain
Coding; Location gains Description, Address and Telecom; Address gains
HouseNumber and GeoPosition.
Value sets: ItemGroupDef.Class was replaced by Class.Name, and SubClass.Name
and SubClass.ParentClass added. The dead Location.LocationType key was removed — the
model attribute has been Role (free text) for some time, so that key matched nothing
and left Role unvalidated. CodeListItem.ExtendedValue was already present and starts
resolving now that the attribute exists.
Parameter, ReturnValue and MethodSignature moved earlier in model.py so that
RangeCheck can reference MethodSignature. Class order in the module carries no
meaning beyond definition-before-use.
Note for anyone extending these classes. Several additions insert a descriptor in the
middle of a class body, which is how child order is expressed. verify_order() reads
each class's own body, so a subclass of an odm_2_0 model class that does not opt into
merge_fields=True must redeclare inherited children in the new order. No shipped model
uses merge_fields, so this affects third-party extensions only.
Added — ODM v2.0 model/XSD alignment (phase 4)
The Protocol study-design subtree, deliberately left unmodelled when Protocol was
first aligned earlier in this release. Twenty-four new classes fill the nine optional
child slots the XSD defines, in XSD sequence order:
StudySummary→StudyParameter→ParameterValue— named summary parameters
such as trial blinding schema.TrialPhase— the trial phase, value-set checked against the 13 XSD terms.StudyIndications→StudyIndication, andStudyInterventions→
StudyIntervention(withStudyInterventionRef).StudyObjectives→StudyObjective, andStudyEndPoints→StudyEndPoint
(withStudyEndPointRef), so an objective can point at the endpoints assessing it.StudyTargetPopulation(withStudyTargetPopulationRef).StudyEstimands→StudyEstimand→IntercurrentEvent/SummaryMeasure— the
ICH E9(R1) estimand framework, which ODM 2.0 models natively.InclusionExclusionCriteria→InclusionCriteria/ExclusionCriteria→
Criterion, each criterion pointing at aConditionDef.
All nine Protocol children are minOccurs="0", so existing documents are unaffected.
Value sets: TrialPhase.Value and StudyEndPoint.Level added. StudyObjective.Leve
— a truncated key that matched nothing, leaving StudyObjective.Level unvalidated — was
corrected to StudyObjective.Level. StudyEndPoint.Type and StudyEstimand.Level were
already present and start resolving now that their classes exist.
With this phase every ODM 2.0 metadata element the XSD defines is modelled. The 24 XSD
elements still without a class are the ClinicalData/ReferenceData data layer, which
ROADMAP.md books for v0.3.0.
Changed — ODM v2.0 model/XSD alignment (phase 5)
The final alignment phase: record what odmlib deliberately does not express, and remove
what the ODM 2.0 XSD does not define. Every model class now corresponds to an ODM 2.0
XSD element, and every metadata element the XSD defines is modelled.
Removed seven ODM 1.3.2 carry-over classes that have no ODM 2.0 XSD element:
ArchiveLayout, Email, ExceptionEvent, Fax, Pager, Phone and Picture
(DisplayName went in phase 1). None was reachable from ODM — no descriptor anywhere
in the model referenced them — so nothing can be lost from a document; they could only
be constructed in isolation and never attached. ODM 2.0 folds Email/Fax/Pager/
Phone into Telecom with a TelecomType of the same name, and replaces Picture
with Image.
Three approximations are now documented rather than latent, each with a
.. note:: Known approximation in its class docstring, in Known Limitations in
README.md, and in the ODM 2.0 section of the model reference guide. In each case
odmlib's descriptor model cannot express what the XSD says, so
ODMSchemaValidator — not object construction — is what catches a violation:
FormalExpression— the XSD requires exactly one ofCodeorExternalCodeLib.
odmlib has no way to express anxs:choice, so both are optional; setting neither, or
both, builds an object odmlib accepts and the schema rejects.odm_2_0has no Cerberus
rules package to enforce it in either.StudyEventGroupDef— the XSD's repeating(StudyEventGroupRef?, StudyEventRef?)
group permits the two to interleave. Two parallel lists emit all groups then all
events. Reading is affected too: the loader collects children by tag, so an interleaved
source document loads correctly and re-serializes grouped. Both forms are schema-valid;
only the ordering is lost.TranslatedText— the XSD types itmixed="true"with an optionalxhtml:div
child, so text may carry XHTML markup. odmlib models the text-only form and drops an
xhtml:divon load. Faithful support would need a model class for every element
ODM-xhtml.xsdallows, since the loader resolves children by tag name; and odmlib
never reads ElementTree'stail, so text following a child element is lost regardless.
Plain-textTranslatedTextround-trips exactly.
model.pyi was brought into line at the same time: the eight stub-only classes naming
elements the model has never had were removed, and CodeList, User, UserName,
GivenName, FamilyName and Image were regenerated or added so that no stub
annotation refers to an undefined class.
Fixed — ODM v2.0 value sets
The odm_2_0 block of odmlib/data/valuesets.json was wrong in both directions. Checking
it against what each attribute's XSD type actually permits — rather than only asking which
keys resolve — turned up eight defects that no test caught.
odmlib rejected schema-valid values in three places. ItemGroupDef.Type and
TrialPhase.Value have XSD types that union an enumeration with bare xs:string, making
them extensible vocabularies, but were enforced as closed lists, so a sponsor-specific
value raised. ODM.ODMVersion is a pattern admitting 2.0.1 and 2.0-draft, stored as
the single literal "2.0".
Two lists held the wrong values. MethodDef.Type accepted Other, which ODM 2.0 does
not define, and rejected Preload, which it does. User.UserType offered four values
where ODM 2.0 has nine, rejecting Subject, Monitor, Data analyst, Care provider
and Assessor. Both had been copied from the odm_1_3_2 block — which is correct for ODM
1.3.2; ODM 2.0 changed both enumerations and the copy never caught up.
Standard was not value-checked at all. Name, Type and PublishingSet are closed
enumerations in the XSD but were modelled as plain strings, so Standard(Type="Nonsense")
built without complaint on a reachable, mostly-required element. They are now
ValueSetString. Standard.Status is an extensible union and takes the open form instead.
Nine keys matching no descriptor were removed, four of them misspellings of attribute names
on ClinicalData classes that arrive in v0.3.0 (AuditRecord.EditPoin,
Comment.SponsorOrSit, Query.SourceSyste, Query.Status). They were not corrected and
kept: a key for a class that does not exist cannot be verified, which is how the
misspellings survived in the first place. v0.3.0 adds them with their classes.
New in odmlib/valueset.py: a third entry form for extensible vocabularies,
{"_values": [...], "_open": true}, alongside the existing list and _regex forms. Any
value is accepted; the listed terms remain the documented ones and drive describe().
Guard. tests/test_odm_2_0_xsd_alignment.py compared value-set key names in both
directions and never looked at values, which is why this survived four alignment phases.
It now classifies each attribute's XSD type as closed enumeration, extensible union,
pattern or free, and reports a missing key, wrong values, a closed list where the schema is
open, or an open entry where the schema is closed. All three value-set allowlists are empty.
A note on the two directions, since several code comments had it backwards: a value-set key
with no descriptor is inert — unused data. A ValueSetString descriptor with no key is
the loud one: validate() maps the unknown sentinel to False and the descriptor then
raises for every value, so the attribute cannot be set at all.