DataFormat

Template

Applies to:  Text Variable

Description

MotifLab provides several data formats to allow users to output their data and analysis results in many different ways. Nevertheless, there may be times when the functionality offered by these formats is not enough, for example if a user has run several analyses and created multiple data objects and wants to combine information from all of these in a single compact document that follows a predefined layout.
The Template format (and its companion TemplateHTML) can be used in such cases to achieve greater control over the presentation of the output. The data formats requires a "template" document which must be provided in the form of a Text Variable. This template can contain regular text which can be interspersed with references to named data objects on the form: {dataobject}. When the template Text Variable is output in Template-format, the references to data objects in the text will be replaced with the actual contents of these data objects. No error message or warning will be given if the data object named in the reference does not exist or if there is some other mistake in the reference. In such cases, MotifLab will simply leave the original reference as is in the output. Note that only "simple" data objects such as Numeric Variables, Text Variables, Collections and OutputData objects can be referenced directly in the template. However, other types of data (such as feature data, partitions or maps) can be included by first outputting the original data object to an intermediate OutputData object in a selected data format. This intermediate object can then be referenced in the template.

The contents of a referenced data object will usually be included directly in a default format, but some control over the presentation is provided with additional options that can be specified after the name of the data object in the reference following a colon: {dataobject:options}.
Some data types can even have multiple options like so: {dataobject:option1:option2:etc...}

Available format-options for different data types:
Data typeOptionDescription
Numeric Variable PrecisionThe number of decimals to include for real values in the output can be specified in the option.
For example, to output the variable X with 2 decimals use the reference code: {X:2}.
PercentageIf the Numeric Variable holds a value between 0 and 1.0 which should be output as a percentage number, simply place a percentage sign in the option. The value of the variable will then automatically be multiplied by 100 and suffixed with a %-sign in the output. This option can be combined with the precision option by writing the %-sign after the number of decimals. If you want a space between the value and the %-sign in the output you can place the space-code "\s" before the %-sign (or if you combine precision number with the %-option you can just place a normal space between them).

Examples: The following shows different output for the variable X=0.2394278.
The reference code {X:%} will output 23.94278%
The reference code {X:\s%} will output 23.94278 %
The reference code {X:0%} will output 23%
The reference code {X:2%} will output 23.94%
The reference code {X:3 %} will output 23.942 %
Collection SeparatorCollections are normally output as sorted lists with entries separated by commas. However, it is also possible to specify a different separator as an option. E.g. to output the collection JASPAR with entries separated by semicolons instead of commas, use the reference code {JASPAR:;}. A space-separator can be specified as "\s", TAB-separator as "\t" and linebreaks as "\n" for plain text or "<br>" for HTML. To use a colon as a separator, simply type "colon", e.g. {JASPAR:colon}.
SizeTo output the size of a collection rather than listing the entries, use the "size" option, or equivalently enclose the name of the collection in vertical bars. E.g. {JASPAR:size} or {|JASPAR|}.
Text Variable

   and

Output Data
SeparatorThe contents of Text Variables and Output Data objects are usually included verbatim in the final document one line at a time. However, it is possible to use a different separator for the lines rather than the usual linebreak separator (which is "\n" for plain text and "<br>" for HTML). This separator is then specified as the first option. Spaces, TABs and colons can be used as separators as described for the Collections type above.
SortingIt is possible to sort the lines from the Text Variable or Output Data according to a natural order by specifying a second option (after the first "separator" option). There are many ways to do this. If the second option includes the text "sort", "asc" or "A" the lines will be sorted in ascending order (e.g. "sort", "sorted", "sort ascending", "ascend","sortA" or simply "A"). However, if the option contains the text "des" or "D" they will be sorted in descending order (e.g. "sort desc", "sort descending", "descend", "desc", "sortD" or simply "D").

The following examples show how the lines from the Text Variable X can be sorted. Note that since sorting must be specified as a second option, the "separator" must be included as the first option (but by just leaving it blank the default separator will be used).
The reference code {X::sort} will output X with lines in ascending order separated by default linebreaks.
The reference code {X::ascending} will output X with lines in ascending order separated by default linebreaks.
The reference code {X::D} will output X with lines in descending order separated by default linebreaks.
The reference code {X:\t:asc} will output X with lines in ascending order separated by TABS.
Remove duplicates The second option can also be used to specify that duplicate lines in the Text Variable or Output Data object should be removed by including the text "uniq" or "U" anywhere in the option. Note that this can be combined with the sorting option, as shown in the last three examples below:

The reference code {X::uniq} will output all unique lines from X separated by default linebreaks.
The reference code {X::sorted unique} will output all unique lines from X sorted in ascending order.
The reference code {X::unique ascending} will output all unique lines from X sorted in ascending order.
The reference code {X:,:DU} will output all unique lines from X in descending order separated with commas.


Example:
In the following scenario, a user has created a protocol script to find significant transcription factor motifs in promoters for sets of sequences that are either up- or downregulated at three different timepoints (1h, 2h or 3h). The sequence collections are named Up1, Up2, Up3, Down1, Down2 and Down3, and the collections of significant motifs are called Motifs_Up1, Motifs_Up2, Motifs_Up3, Motifs_Down1, Motifs_Down2 and Motifs_Down3. What the user wants now is to create a simple report that contains information about how many sequences there were in each such collection and list the names of the significant TFs for each collection. The following protocol first creates six OutputData objects containing names of significant TFs and then creates the final report based on a predefined template.
Motifs_Up1_out = output Motifs_Up1 in Motif_Properties format {Format="Clean Short name"}
Motifs_Up2_out = output Motifs_Up2 in Motif_Properties format {Format="Clean Short name"}
Motifs_Up3_out = output Motifs_Up3 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down1_out = output Motifs_Down1 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down2_out = output Motifs_Down2 in Motif_Properties format {Format="Clean Short name"}
Motifs_Down3_out = output Motifs_Down3 in Motif_Properties format {Format="Clean Short name"}

TemplateText = new Text Variable(File:"Template_text.txt", format=Plain)
Output1= output TemplateText in Template format

The file "Template_text.txt" which is used in the protocol above to provide the template text for the final report is shown below. All references to data objects in the template are marked in red color. First the template produces a 2x3-table with the sizes of each of the sequence collections. Then it writes out 6 lines listing the significant transcription factors for each such collection. It would be possible to reference the motif collections directly in the template, e.g. {Motifs_Up1}. However, this would then only list the IDs of the motifs and not the TF-names, so instead the protocol above uses the Motif_Properties data format to output the "clean short name" of each motif in each collection to an intermediate OutputData object that is referenced instead (Motifs_Up1_out). These OutputData objects contains one TF-name on each line and there might be duplicate names if there are several motif models for the same TF. Hence, the template specifies that the names in the OutputData objects should be sorted alphabetically and duplicate names should be removed (using the "sorted unique" option). Also, rather than using a linebreak to separate the TF-names, a comma should be used instead.
Number of genes up- and down-regulated at different time points:

Time    1h       2h       3h
--------------------------------
UP      {Up1:size}    {Up2:size}   {Up3:size}
DOWN    {Down1:size}  {Down2:size} {Down3:size}
================================

Significant transcription factors in promoters of these genes:

Up 1h: {Motifs_Up1_out:,:sorted unique}
Down 1h: {Motifs_Down1_out:,:sorted unique}
Up 2h: {Motifs_Up2_out:,:sorted unique}
Down 2h: {Motifs_Down2_out:,:sorted unique}
Up 3h: {Motifs_Up3_out:,:sorted unique}
Down 3h: {Motifs_Down3_out:,:sorted unique}

The result could look something like this:
Number of genes up- and down-regulated at different time points:

Time      1h     2h     3h
--------------------------------
UP        23     42     31
DOWN      17     35     26
================================

Significant transcription factors in promoters of these genes:

Up 1h: EGR1,FOS1
Down 1h: 
Up 2h: FOXD1,GATA2,HOXA5,IRX5,JUND
Down 2h: E2F5,KLF15,NR2C1
Up 3h: ATF5,NFX1,NKX3-1,ZHX2
Down 3h: EVI1,HEYL,HOXC9,POU2F1

See Also: TemplateHTML, output, Text Variable, Output Data