Skip to content

How to handle reads with mixed primer sequences #875

Description

@tao-bioinfo

This is somewhat similar to #744 , but they are totally different.

An example is https://www.ncbi.nlm.nih.gov/sra/SRR31443117, containing 173,536 reads. It is a mixture of two primer sets targeting adjacent but different regions.

The first primer set is:

Command line parameters: -a GTCGGTAAAACTCGTGCCAGC;required...CAAACTGGGATTAGATACCCCACTATG;optional --no-indels -e 4.5 

Output:

Total reads processed:                 173,536
Reads with adapters:                    44,695 (25.8%)

The second primer set is:

Command line parameters: -a ACTGGGATTAGATACCCC;required...CTAGAGGAGCCTGTTCTA;optional --no-indels -e 4.5 

Output:

Total reads processed:                 173,536
Reads with adapters:                   144,909 (83.5%)

How to deal with such situation efficiently? Run with two steps and combine them seems does not appropriate, since there is a small overlap between these two groups (25.8% + 83.5% = 109.3% > 100% )

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions