Process Data


During upload, Data and Phenotype Files in the expected format will be processed for tagging and summary statistics will be calculated.

Tagging

Uploaded data files are processed for tagging. This tagging process expands the search space to include the contents in the original files, molecule synonyms, and related identifiers. Tagging will only work on recognized data types; however, you can contact the mapMECFS team to request a new data type.

Recognized Data Types for Tagging

Tagging allows users to search the contents in the original source file as well as an expanded search of relevant databases (see table below). The recognized data types for tagging include:

  • Cytokine Assay
  • Demographic, Health, and Survey (DHS)
  • Gene Expression
  • Metabolomics
  • Methylation
  • miRNA
  • Proteomics

Note: Upload is not restricted to these data types; any data files can be uploaded to mapMECFS; tags will be cleaned annually by the Data Management and Coordinating Center for consistency across studies and to eliminate redundancy.

Data Type Required Data Column(s) Database used for Tagging What is Searchable? Example Searches
Cytokine Assay
  • Molecule
NCBI Gene (December 2021)
  • Any entry from the Molecule data column
  • Matching gene synonyms
Demographic, Health, and Survey
  • ParticipantID
  • Phenotype
N/A
  • Any entry from the ParticipantID data column
Gene Expression
  • Molecule
NCBI Gene (December 2021)
  • Any entry from the Molecule data column
  • Matching gene synonyms
Metabolomics
  • InChiKey
  • Molecule
  • database_identifier
N/A
  • User input from any of the three required columns
Methylation
  • Molecule
Illumina 450K (v.15017482_v1-2) or Infinium MethylationEPIC (v-1-0-b4). Please email mapmecfs@rti.org if another manifest file is needed.
  • Any entry from the Molecule data column
  • Corresponding B37 coordinates (Chr:Pos)
miRNA
  • Molecule
miRBase (March 2019)
  • Any entry from the Molecule data column
  • Any miRNA related to the primary transcript
  • Any matching alias
Other
  • Molecule
N/A
  • Any entry from the Molecule data column
Proteomics
  • Molecule
NCBI Gene (December 2021)
  • Any entry from the Molecule data column
  • Matching protein/gene synonyms

Summary Statistics

For recognized Data Types mapMECFS generates a Summary Statistics file to characterize how dataset measures compare between phenotype groups as annotated in the uploaded Phenotype File. A nonparametric Wilcoxon rank-sum test is used to distinguish how dataset features differ between groups within the study (e.g., between cases and controls). Summary statistics are automatically calculated for each feature in the uploaded gene expression, cytokine assay, metabolomics, miRNA, or methylation Data Files when a correctly formatted Phenotype File is uploaded to mapMECFS.

Please note that summary statistics are processed asynchronously to avoid impacting load times. Therefore, they may not be immediately available after upload as the calculations are made in the background.

Once calculations are complete, one can view the resulting Summary Stats file by opening the dataset of interest and scrolling to Summary Statistics < View Summary Statistics.

Summary columns in this file include:

  • Sample sizes in each group (labeled as "count")
  • Median value for each group
  • Standard deviation
  • Wilcoxon rank-sum test statistic (labeled as "Ranksum stat")
  • Wilcoxon rank-sum p-value (labeled as "Ranksum p-value")
  • Wilcoxon rank-sum Bonferroni Corrected p-value (labeled as "Ranksum Bonf")