• FluTrackers.com Inc. does not provide medical advice. Information on this web site is collected from various internet resources, and the FluTrackers board of directors makes no warranty to the safety, efficacy, correctness or completeness of the information posted on this site by any author or poster. The information collated here is for instructional and/or discussion purposes only and is NOT intended to diagnose or treat any disease, illness, or other medical condition. Every individual reader or poster should seek advice from their personal physician/healthcare practitioner before considering or using any interventions that are discussed on this website. By continuing to access this website you agree to consult your personal physican before using any interventions posted on this website, and you agree to hold harmless FluTrackers.com Inc., the board of directors, the members, and all authors and posters for any effects from use of any medication, supplement, vitamin or other substance, device, intervention, etc. mentioned in posts on this website, or other internet venues referenced in posts on this website.
  • We are not asking for any donations. Do not donate to any entity who says they are raising funds for us.

Identifying mutation positions in all segments of influenza genome enables better differentiation between pandemic and seasonal strains

tetano

Editor, Senior Moderator
Gene. 2019 Feb 12. pii: S0378-1119(19)30058-7. doi: 10.1016/j.gene.2019.01.014. [Epub ahead of print]
[h=1]Identifying mutation positions in all segments of influenza genome enables better differentiation between pandemic and seasonal strains.[/h] Kargarfard F[SUP]1[/SUP], Sami A[SUP]2[/SUP], Hemmatzadeh F[SUP]3[/SUP], Ebrahimie E[SUP]4[/SUP].
[h=3]Author information[/h]

[h=3]Abstract[/h] Influenza has a negative sense, single-stranded, segmented RNA. In the context of pandemic influenza research, most studies have focused on variations in the surface proteins (Hemagglutinin and Neuraminidase). However, new findings suggest that all internal and external proteins of influenza viruses can contribute in pandemic emergence, pathogenicity and increasing host range. The occurrence of the 2009 influenza pandemic and the availability of many external and internal segments of pandemic and non-pandemic sequences offer a unique opportunity to evaluate the performance of machine learning models in discrimination of pandemic from seasonal sequences using mutation positions in all segments. In this study, we hypothesized that identifying mutation positions in all segments (proteins) encoded by the influenza genome would enable pandemic and seasonal strains to be more reliably distinguished. In a large scale study, we applied a range of data mining techniques to all segments of influenza for rule discovery and discrimination of pandemic from seasonal strains. CBA (classification based on association rule mining), Ripper and Decision tree algorithms were utilized to extract association rules among mutations. CBA outperformed the other models. Our approach could discriminate pandemic sequences from seasonal ones with more than 95% accuracy for PA and NP, 99.33% accuracy for NA and 100% accuracy, precision, specificity and sensitivity (recall) for M1, M2, PB1, NS1, and NS2. The values of precision, specificity, and sensitivity were more than 90% for other segments except PB2. If sequences of all segments of one strain were available, the accuracy of discrimination of pandemic strains was 100%. General rules extracted by rule base classification approaches, such as M1-V147I, NP-N334H, NS1-V112I, and PB1-L364I, were able to detect pandemic sequences with high accuracy. We observed that mutations on internal proteins of influenza can contribute in distinguishing the pandemic viruses, similar to the external ones.
Copyright ? 2019 Elsevier Inc. All rights reserved.


[h=4]KEYWORDS:[/h] Association rule mining; CBA; Expert system; Hot spots; Pandemic influenza; Ripper algorithm

PMID: 30769139 DOI: 10.1016/j.gene.2019.01.014
 
What happened to HA - accuracy values for the other RNA strands are given? Also I was surprised by PB2 being an outlier (although it does not state a %) as I would expect it to be closely correlated to PB1,PA and NP as they form RNP and work as a biological functional unit.[1]
I was very pleased to see the way they used the growing sequence data set as it is beginning to use it as I argued in my essay on the need for changes in sequence collection. [2]

[1] The end of post #50 plus post #79 looks at these strands and the RNP.
https://flutrackers.com/forum/forum...l-1-2013-to-june-3-2013-closed/page4?t=202715

[2] The first post argues that the sequence databases are not being built optimally the second is more relevant focusing on how it could be data mined if more care was taken in collecting the right samples.
https://flutrackers.com/forum/forum...fluenza-databases-need-root-and-branch-reform
 
Back
Top Bottom