A finished graphic tells a model what a design looks like. An editable source can also expose how the design is constructed: text objects, image datasets, vectors, hierarchy, geometry, grouping, masks, transformations, and relationships between components. That difference matters when the target system must do more than generate pixels. It may need to reconstruct a design, change one element without disturbing another, adapt a layout to a new format, or return an artifact that a human activity recognition datasets can continue editing.
But “more structure” is not the same as “better training data.” A source file can be mostly flattened, contain missing fonts or linked assets, use application - specific features that disappear during parsing, or encode document organization that is mistaken for semantic truth. It also costs more to process, validate, govern, and sometimes license. The useful question is therefore narrower: does preserving source structure improve the capability the model is supposed to deliver?
This guide focuses on that question. It separates rendered images, native source structure, semantic annotations, paired design states, and operation traces; explains how they should be represented and validated; and shows how AI teams can test whether editable source data provides measurable value.
Table of contents:
- ● What Is Editable Design Training Data?
- ● Flat Images and Editable Source Files Represent Different Information
- ● Generation, Reconstruction, and Editing Are Different Model Problems
- ● What Source Structure Actually Adds
- ● Source - to - Render Pairs: Useful Only When the Link Is Reproducible
- ● Normalization Can Improve Consistency and Still Lose Information
- ● Which Source Formats Are Useful?
- ● What Recent Research Shows
What Is Editable Design Training Data?
Editable design training data preserves some of a creative artifact’s construction instead of storing only its final rasterized appearance. Depending on the source format and extraction process, the dataset may retain:
• Layers and groups
• Objects and components
• Editable text and typography properties
• Vector paths and geometry
• Masks, clipping, effects, and blending - related properties
• Coordinates, dimensions, transforms, and stacking order
• Linked or embedded assets
• Component or document relationships
• Version information and related design states
The extension is not the important part. A PSD can contain rich layer structure or be largely flattened. An SVG can expose explicit vector objects and transforms. A design - tool API may expose nodes, components, and relationships that are absent from a rendered export. A useful dataset therefore measures the structure that survives extraction rather than inferring utility from the file name.

For an AI pipeline, four representations are especially useful:
Representation | What it contributes |
|---|---|
Render | Finished visual appearance and pixel-level evidence. |
Source structure | Components, hierarchy, geometry, editability, and native properties. |
Semantic annotation | What an element represents, such as headline, logo, or CTA. |
Related states / operation traces | What changed between states, and where available, how the change was produced. |
The target output should drive the data design. A product that only needs a final JPEG has a different information requirement from a product that must return a structured, editable document.
Flat Images and Editable Source Files Represent Different Information
A promotional banner might contain a product image, headline, supporting copy, logo, call to action, decorative elements, and background. The rendered image shows their visible arrangement. It does not directly state which elements are editable text, which objects are grouped, which shapes are vector paths, or which spacing relationships come from reusable layout structure.
Those properties can sometimes be inferred from pixels. The distinction is that inference is a model - generated hypothesis, while source structure can provide the underlying representation itself. For tasks that depend on object identity, hierarchy, or editability, that difference can be useful supervision.
Capability | Flat image | Editable source |
|---|---|---|
Final appearance | Directly represented | Available through a render |
Object separation | Usually inferred | Often explicit |
Layer hierarchy | Absent | Often preserved |
Editable text | Not preserved | Often preserved |
Vector geometry | Rasterized | May be preserved |
Component relationships | Mostly inferred | May be represented |
Processing burden | Lower | Higher |
Structural supervision | Limited | Potentially substantial |
Generation, Reconstruction, and Editing Are Different Model Problems
“Prompt to finished image” and “prompt to structured editable design” should not be treated as the same target. The second requires the system to generate or preserve objects, hierarchy, text, geometry, relationships, and editable components. The training data has to expose enough information for those outputs to be learned and evaluated.
The same applies to editing. An instruction such as “replace the headline while preserving the logo, product image, and geometry” is not simply a smaller text - to - image task. The system has to change one part and preserve specified parts. A dataset that includes related source states or explicit edit instructions can represent that constraint more directly than a collection of unrelated final images.
What Source Structure Actually Adds

Structure and relationships
Source documents can make component identity and relationships explicit. Useful fields include object type, group membership, coordinates, dimensions, alignment, containment, stacking order, and relative scale. These fields are especially relevant to layout generation because design quality depends on spatial relationships as well as appearance.
Structure still has limits. A layer named “Layer 23” does not establish that it is a hero product. A group called “Header” does not guarantee a semantic role. A source file describes how a document is organized; it does not automatically describe the designer’s intent, audience, or business purpose.
Editability is not semantic labeling
Treat native source properties, derived fields, and semantic labels as different evidence classes. For example:
Provenance class | Example |
|---|---|
Native source property | Text string, layer coordinates, native object type |
Derived property | Parser-inferred object category or normalized relationship |
Semantic annotation | Primary CTA, hero image, intended audience, visual style |
Field - level provenance should travel with the record where practical. A coordinate extracted by a parser should not have the same evidence status as a “primary CTA” label created by a human reviewer. Distinguishing deterministic, extracted, model - generated, and human - reviewed fields makes the dataset easier to audit and easier to improve.
Paired States Are Not Operation Traces
Related versions can provide transformation supervision, but only to the extent that the relationship between versions is actually known.
Data | What it tells you |
|---|---|
Single source state | What the design contains at one point in time |
Paired states | What differs between two documented states |
Operation trace | Which operations produced the change, when that history is explicitly recorded |
If Design A and Design B differ, the pair does not necessarily tell you which instruction caused the change, what operations were performed, their order, whether the edit was automatic or manual, or which elements were intended to stay fixed. A before - and - after pair describes a transformation outcome. An operation trace describes the transformation process.
Source - to - Render Pairs: Useful Only When the Link Is Reproducible

A source file answers “what can be edited?” A render answers “what did the composition look like under a particular rendering environment?” Keeping both creates a useful connection between structure and appearance, but the relationship should not be modeled as a single linear pipeline.
A more accurate architecture is:
Source document - > parser - > structured representation
Source document - > renderer - > reference render
Metadata, semantic annotations, dependency records, and quality evidence can attach to either representation or to the relationship between them.
Why Object - to - Pixel Mapping Is Not Automatic
A source object does not always correspond to one isolated visible region. Group opacity, masks, clipping paths, blend modes, adjustment layers, filters, nested transforms, and neighboring objects can change the final appearance. If object - level visual attribution matters, useful supplementary records can include object masks, isolated layer renders, visibility maps, stable element IDs, or explicit source - to - render correspondence annotations.
The distinction is important for training and evaluation: source hierarchy and pixel - level attribution are related representations, not interchangeable ones.
Renderer and Dependency Reproducibility
A reference render should be generated under a documented environment and retained as evidence of the expected output. At minimum, the record should identify the relevant source version, rendering application or engine, renderer version, fonts, external or embedded assets, plugin/effect dependencies, color profile, output dimensions, and a stable render identifier or hash.
Validation item | Question |
|---|---|
Source validity | Does the document open or parse correctly? |
Dependency validity | Are fonts, linked assets, plugins, and other requirements available? |
Structural validity | Is the required editable structure present? |
Render validity | Does the source render as expected? |
Semantic validity | Are added labels or roles correct? |
Task validity | Does the representation help the intended model task? |
These are separate failure modes. A file can parse correctly and still render incorrectly. It can render acceptably while losing the structure the model needs. A visually plausible result can also hide a missing font or asset because a fallback was substituted.
Normalization Can Improve Consistency and Still Lose Information

A common schema makes mixed - source datasets easier to train and integrate, but normalization is a representation decision, not a lossless administrative step. PSD, SVG, Figma, InDesign, and other design systems expose different concepts and capabilities. A common schema can simplify or discard application - specific effects, typography features, component semantics, masks, layout constraints, or reusable style references.
A practical architecture keeps the raw source, a normalized common representation, and source - specific extensions. That gives downstream teams a consistent core without forcing every source system into the same abstraction.
File Count Is Not the Same as Design Coverage
A large editable archive can contain many files derived from a much smaller set of underlying design families. Alternate aspect ratios, localizations, color variants, product substitutions, campaign revisions, exports, and saved snapshots may be useful training examples, but they are not all independent design concepts.
Report file count separately from independent design - family count and transformation - pair count. This distinction matters for both dataset analysis and evaluation because closely related files can make a collection appear more diverse than it really is.
Template - Family Leakage Can Distort Evaluation
Suppose training contains a desktop version of a template and evaluation contains a mobile version of that same template. The evaluation file may be new, but the underlying design family is familiar. If the claim is generalization to unseen templates, the family should be grouped before splitting.
Generalization claim | Recommended split unit |
|---|---|
Unseen file variants | File |
Unseen template families | Template family |
Unseen campaigns | Campaign |
Unseen layout families | Layout family |
Unseen source projects | Parent/source project |
The split unit should match the generalization claim. There is no universal “best split”; the question is what kind of novelty the evaluation is supposed to measure.
Which Source Formats Are Useful?

Native application files, vector formats, design - tool APIs, and programmatic representations can all support structured design training. What matters is the information that can be extracted reliably, not the prestige of the format.
A PSD may preserve layers, text, masks, effects, and Smart Objects, but those features can introduce dependencies and proprietary semantics. SVG provides explicit vector structure. Design - tool APIs can expose nodes and components. Programmatic representations such as HTML/CSS can encode editability without requiring a native creative - tool file. The representation should be selected around the target capability and the fidelity of the extraction pipeline.
What Recent Research Shows

Jen Togonon
