Description
MotifLab provides several data formats to allow users to output their data and
analysis results in many different ways. Nevertheless, there may be
times when the functionality offered by these formats is not enough,
for example if a user has run several analyses and created multiple data objects and wants to combine
information from all of these in a single compact document with follows a
predefined layout.
The
TemplateHTML format (and its companion
Template) can be
used in such cases to achieve greater control over the presentation of the
output. The data format requires a "template" document which must be provided in
the form of a Text Variable. This template can contain regular text and
HTML-markup which can be interspersed with references to named data objects on the form:
{dataobject}. When the template Text
Variable is output in TemplateHTML-format, the references to data objects in the
text will be replaced with the actual contents of these data objects.
No error message or warning will be given if the data object named
in the reference does not exist or if there is some other mistake in the
reference. In such cases, MotifLab will simply leave the original
reference
as is in the output.
Note that only "simple" data objects such as Numeric Variables, Text Variables, Collections and OutputData
objects can be referenced in the template. However, other types of data
objects (such as feature data, partitions or maps) can be included by first
outputting these to intermediate OutputData objects which can then be
referenced.
The contents of a referenced data object will usually be included directly in a
default format, but some control over the presentation is provided with additional
options which are described in the documentation for
the
Template format.
Example:
In the following scenario, a user has created a protocol script to find
significant transcription factor motifs in promoters for sets of sequences
that are either up- or downregulated at three different timepoints (1h, 2h or
3h). The sequence collections are named Up1, Up2, Up3, Down1, Down2 and Down3,
and the collections of significant motifs are called Motifs_Up1, Motifs_Up2,
Motifs_Up3, Motifs_Down1, Motifs_Down2 and Motifs_Down3. What the user wants
now is to create a simple report that contains information about how many
sequences there were in each such collection and list the names of the
significant TFs for each collection. The following protocol first creates six
OutputData objects containing names of significant TFs and then creates
the final report based on a predefined template.
Motifs_Up1_out = output Motifs_Up1 in Motif_Properties format {Format="Clean Short name"}
Motifs_Up2_out = output Motifs_Up2 in Motif_Properties format {Format="Clean Short name"}
Motifs_Up3_out = output Motifs_Up3 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down1_out = output Motifs_Down1 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down2_out = output Motifs_Down2 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down3_out = output Motifs_Down3 in Motif_Properties format {Format="Clean Short name"}
TemplateText = new Text Variable(File:"Template_text.html", format=Plain)
Output1= output TemplateText in TemplateHTML format
The file "Template_text.html" which is used in the protocol above to provide
the template text for the final report is shown below.
All references to data objects in the template are marked in red color.
First the template produces a 2x3-table with the sizes of each of the sequence collections. Then it
writes out a second table listing the significant transcription factors for each such
collection. It would be possible to reference the motif collections
directly in the template, e.g. {Motifs_Up1}. However, this would then only list
the IDs of the motifs and not the TF-names, so instead the protocol above uses
the Motif_Properties data format to output the "clean short name" of each
motif in each collection to an intermediate OutputData object that is referenced instead (Motifs_Up1_out).
These OutputData objects contains one TF-name on each line and there might be
duplicate names if there are several motif models for the same TF. Hence, the
template specifies that the names in the OutputData objects should be sorted
alphabetically and duplicate names should be removed (using the "AU" option).
<h2>Number of genes up- and down-regulated at different time points:</h2>
<table>
<tr><th>Time</th><th>1h</th><th>2h</th><th>3h</th></tr>
<tr><td>Up</td><td>{Up1:size}</td><td>{Up2:size}</td><td>{Up3:size}</td></tr>
<tr><td>Down</td><td>{Down1:size}</td><td>{Down2:size}</td><td>{Down3:size}</td></tr>
</table>
<br>
<h2>Significant transcription factors in promoters of these genes:</h2>
<table>
<tr>
<th style="background-color:#E0E0E0;">1h</th>
<th style="background-color:#C8C8C8;">2h</th>
<th style="background-color:#B0B0B0;">3h</th>
</tr>
<tr>
<td valign=top style="background-color:#FFD0D0;">{Motifs_Up1_out::AU}</td>
<td valign=top style="background-color:#FFC0C0;">{Motifs_Up2_out::AU}</td>
<td valign=top style="background-color:#FFC0B0;">{Motifs_Up3_out::AU}</td>
</tr>
<tr>
<td valign=top style="background-color:#D0FFD0;">{Motifs_Down1_out::AU}</td>
<td valign=top style="background-color:#C0FFC0;">{Motifs_Down2_out::AU}</td>
<td valign=top style="background-color:#B0FFB0;">{Motifs_Down3_out::AU}</td>
</tr>
</table>
The result could look something like this:
Number of genes up- and down-regulated at different time points:
| Time | 1h | 2h | 3h |
| Up | 23 | 42 | 31 |
| Down | 17 | 35 | 26 |
Significant transcription factors in promoters of these genes:
| 1h |
2h |
3h |
EGR1 FOS1 |
FOXD1 GATA2 HOXA5 IRX5 JUND |
ATF5 NFX1 NKX3-1 ZHX2 |
|
E2F5 KLF15 NR2C1 |
EVI1 HEYL HOXC9 POU2F1 |