{
  "_schema": "https://data.nist.gov/od/dm/nerdm-schema/v0.7#",
  "@context": [
    "https://data.nist.gov/od/dm/nerdm-pub-context.jsonld",
    {
      "@base": "ark:/88434/mds2-4088"
    }
  ],
  "@type": [
    "nrdp:DataPublication",
    "nrdp:PublicDataResource",
    "dcat:Dataset"
  ],
  "_extensionSchemas": [
    "https://data.nist.gov/od/dm/nerdm-schema/pub/v0.7#/definitions/PublicDataResource"
  ],
  "@id": "ark:/88434/mds2-4088",
  "ediid": "ark:/88434/mds2-4088",
  "version": "1.0.0",
  "doi": "doi:10.18434/mds2-4088",
  "title": "Supporting data for \"DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats\"",
  "contactPoint": {
    "fn": "Kevin Kiesler",
    "hasEmail": "mailto:kevin.kiesler@nist.gov"
  },
  "modified": "2026-02-19",
  "status": "available",
  "landingPage": "https://data.nist.gov/od/id/mds2-4088",
  "description": [
    "Polymerase chain reaction (PCR) applications including sequencing rely on accurate thermostable DNA polymerases. Polymerization errors may hinder the detection of low-level DNA variants such as mutations in clinical samples or DNA from minor contributors in crime scene traces. Short Tandem Repeat (STR) markers are particularly affected by artefacts. Apart from the regular random base substitutions, the repeated structure of STRs makes them prone to formation of stutter products. However, the mechanisms leading to stutter formation have not yet been fully elucidated. Here, we applied an STR assay based on Unique Molecular Identifiers (UMIs) to study the effects of DNA polymerases with different characteristics on the amplicon yield as well as the formation of PCR errors. The application of UMIs made it possible to study the impact on error formation of applying genomic DNA (mimicking the early PCR cycles) or amplicons (later cycles) as template. The levels of base substitutions were clearly connected to the fidelity of the DNA polymerases, which in turn was coupled with having an integrated 3\u2019to 5\u2019 exonuclease domain. Stutter formation, on the other hand, was not as directly associated with fidelity, as two high-fidelity polymerases showed quite different levels of stutter. DNA binding domains generally improve processivity which could lower the incidence of stutter. However, this was not clear in the present study as a polymerase having a DNA binding domain gave the highest stutter levels. Overall, the degree of polymerase stuttering is likely due to several different DNA polymerase characteristics. Identifying a DNA polymerase that provides low levels of stutters and base substitutions may enable the detection of low-level variants such as DNA from minor contributors in mixed forensic traces.",
    "The repository contains files from the sequencing instrument, in fastq format, constituting all data from the experiments performed. \nSequence data is included in a single compressed folder, with subfolders containing fastq files generated by a single sequencing run. \nIn total, five sequencing runs were performed to generate all data consisting of 416 fastq files derived from 208 samples.\nThere are five subfolders, named as following:\n250208_reseq_fastq (80 files from 40 samples)\n250221_fastq (96 files from 48 samples)\n250227_fastq (80 files from 40 samples)\nu17pols (80 files from 40 samples)\nu18pols (80 files from 40 samples)",
    "File naming structure:\n\"Number\"-\"Polymerase used in barcoding PCR\"-\"Polymerase used in adaptor PCR\"-\"DNA sample\"-\"Input amount\"-\"replicate\"_S*_L001_R*_001.fastq",
    "Where:\n\"Number\" is the sample barcode I.D. (1 through 88) for each individual library preparation (Note: barcodes 89 through 96 were not used in these experiments)\n\"Polymerase used in barcoding PCR\" is the enzyme used in the first steps of PCR to introduce the Universal Molecular Index sequence tag\n\"Polymerase used in adaptor PCR\" is the enzyme used in the second phase of amplification to generate high concentration PCR products with sequencing adaptors at 5' and 3' ends\n\"DNA sample\" is the name of the DNA template used, which may be: 2800M, NIST SRM 2391d Component C, or one of ten single source samples used as a testbed (SS1 through SS10)\n\"Input amount\" is the quantity of DNA (in ng) used as template for the PCR amplification\n\"replicate\" is the number corresponding to replicate library preparations for the sample (either 1 or 2)\n\"S*\" is sample number in the sequencing run\n\"L001\" is appended by the sequencing instrument indicating Lane001\n\"R*\" is R1 for Read 1 or R2 for Read 2\n\"001\" is a file suffix appended by the sequencing instrument which is always 001"
  ],
  "keyword": [
    "base substitutions",
    "fidelity",
    "in vitro polymerization",
    "PCR",
    "processivity",
    "STR",
    "stutter"
  ],
  "topic": [
    {
      "@type": "Concept",
      "scheme": "https://data.nist.gov/od/dm/nist-themes/v1.1",
      "tag": "Forensics: DNA and biological evidence"
    }
  ],
  "accessLevel": "public",
  "license": "https://www.nist.gov/open/license",
  "publisher": {
    "name": "National Institute of Standards and Technology",
    "@type": "org:Organization"
  },
  "language": [
    "en"
  ],
  "bureauCode": [
    "006:55"
  ],
  "programCode": [
    "006:052"
  ],
  "_editStatus": "done",
  "theme": [
    "Forensics: DNA and biological evidence"
  ],
  "components": [
    {
      "@id": "cmps/Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip",
      "@type": [
        "nrdp:DataFile",
        "nrdp:DownloadableFile",
        "dcat:Distribution"
      ],
      "_extensionSchemas": [
        "https://data.nist.gov/od/dm/nerdm-schema/pub/v0.7#/definitions/DataFile"
      ],
      "filepath": "Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip",
      "downloadURL": "https://data.nist.gov/od/ds/ark:/88434/mds2-4088/Supporting%20data%20for%20DNA%20polymerase%20characteristics%20influence%20noise%20levels%20in%20sequencing%20of%20short%20tandem%20repeats.zip",
      "mediaType": "application/zip",
      "size": 17763598116,
      "checksum": {
        "hash": "431e75b8cfad9c90da5758c8fe27ee7082f6a6e7fa8021654cf9c4569c2b4ab0",
        "algorithm": {
          "tag": "sha256",
          "@type": "Thing"
        }
      }
    },
    {
      "@id": "cmps/4088_README.txt",
      "@type": [
        "nrdp:DataFile",
        "nrdp:DownloadableFile",
        "dcat:Distribution"
      ],
      "_extensionSchemas": [
        "https://data.nist.gov/od/dm/nerdm-schema/pub/v0.7#/definitions/DataFile"
      ],
      "filepath": "4088_README.txt",
      "downloadURL": "https://data.nist.gov/od/ds/mds2-4088/4088_README.txt",
      "mediaType": "text/plain",
      "format": {
        "description": "Plain text"
      },
      "description": "Text file containing a description of the data set.",
      "title": "4088_README.txt",
      "size": 5318,
      "checksum": {
        "hash": "5d5ac3f06227dc8bb9ed90f67c7039306d352cc1acb487fd7f06ea2b09ef25a9",
        "algorithm": {
          "tag": "sha256",
          "@type": "Thing"
        }
      }
    },
    {
      "@id": "cmps/Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip.sha256",
      "@type": [
        "nrdp:ChecksumFile",
        "nrdp:DownloadableFile",
        "dcat:Distribution"
      ],
      "filepath": "Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip.sha256",
      "downloadURL": "https://data.nist.gov/od/ds/ark:/88434/mds2-4088/Supporting%20data%20for%20DNA%20polymerase%20characteristics%20influence%20noise%20levels%20in%20sequencing%20of%20short%20tandem%20repeats.zip.sha256",
      "algorithm": {
        "tag": "sha256",
        "@type": "Thing"
      },
      "describes": "cmps/Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip",
      "description": "SHA-256 checksum value for Supporting data for DNA polymerase characteristics influence noise levels in sequencing of short tandem repeats.zip",
      "_extensionSchemas": [
        "https://data.nist.gov/od/dm/nerdm-schema/pub/v0.7#/definitions/ChecksumFile"
      ],
      "mediaType": "text/plain",
      "size": 64,
      "checksum": {
        "hash": "68166bc208e749c8630514aab6baa51032cf4dbabece927fcca108e29e67df2a",
        "algorithm": {
          "tag": "sha256",
          "@type": "Thing"
        }
      },
      "valid": true
    }
  ],
  "authors": [
    {
      "familyName": "Lindh",
      "fn": "Tova  Lindh",
      "givenName": "Tova",
      "middleName": "",
      "affiliation": [
        {
          "title": "Lund University, Division of Biotechnology and Applied Microbiology",
          "subunits": [
            "Department of Process and Life Science Engineering"
          ],
          "@type": "org:Organization"
        }
      ],
      "orcid": "",
      "@type": "foaf:Person"
    },
    {
      "familyName": "Sidstedt",
      "fn": "Maja  Sidstedt",
      "givenName": "Maja",
      "middleName": "",
      "affiliation": [
        {
          "title": "National Forensic Center",
          "subunits": [
            "Swedish Police Authority"
          ],
          "@type": "org:Organization"
        }
      ],
      "orcid": "",
      "@type": "foaf:Person"
    },
    {
      "familyName": "Kiesler",
      "fn": "Kevin M. Kiesler",
      "givenName": "Kevin",
      "middleName": "M.",
      "affiliation": [
        {
          "title": "National Institute of Standards and Technology",
          "subunits": [
            "Biomolecular Measurement Division"
          ],
          "@type": "org:Organization",
          "@id": "ror:05xpvk416"
        }
      ],
      "orcid": "0000-0001-7995-6328",
      "@type": "foaf:Person"
    },
    {
      "familyName": "Vallone",
      "fn": "Peter M. Vallone",
      "givenName": "Peter",
      "middleName": "M.",
      "affiliation": [
        {
          "title": "National Institute of Standards and Technology",
          "subunits": [
            "Biomolecular Measurement Division"
          ],
          "@type": "org:Organization",
          "@id": "ror:05xpvk416"
        }
      ],
      "orcid": "0000-0002-8019-6204",
      "@type": "foaf:Person"
    },
    {
      "familyName": "Hedman",
      "fn": "Johannes  Hedman",
      "givenName": "Johannes",
      "middleName": "",
      "affiliation": [
        {
          "title": "Lund University, Division of Biotechnology and Applied Microbiology",
          "subunits": [
            "Department of Process and Life Science Engineering"
          ],
          "@type": "org:Organization"
        }
      ],
      "orcid": "",
      "@type": "foaf:Person"
    }
  ],
  "annotated": "2026-04-07T19:18:02.186879",
  "revised": "2026-04-07T19:18:02.186879",
  "issued": null,
  "firstIssued": "2026-04-07T19:18:02.186879"
}