Showing posts with label QSLiMFinder. Show all posts
Showing posts with label QSLiMFinder. Show all posts

Tuesday, 3 December 2013

File management for large SLiMSuite runs

The latest release of SLiMSuite features a slight modification to the way that files are generated and tidied, which can be beneficial for large runs.

Previously, a different results directory (resdir=PATH) was required for each different run to avoid dataset-specific results being over-written. The partial exception was the *.pickle.gz file, which included some SLiMBuild information in its name. (This is predominantly to speed up the ability of (Q)SLiMFinder to recognise when an intermediate pickle file can be used or not.) As of the latest release, the RunID (runid=X) is also now included in dataset-specific output, allowing results from several different runs (with different RunIDs) to go into the same results directory.

The exception is the files that are created as part of the initial setup/SLiMBuild process: *.slimdb, *.dis.tdt and *.upc. From a given Dataset and RunID, the following files will therefore be generated in ResDir/

Dataset.RunID.cloud.txt
Dataset.RunID.mapping.fas
Dataset.RunID.maskaln.fas
Dataset.RunID.masked.fas
Dataset.RunID.motifaln.fas
Dataset.RunID.occ.csv
Dataset.dis.tdt
Dataset.#SLiMBuild-Text#.pickle.gz
Dataset.slimdb
Dataset.upc

Note that the default ResDir is SLiMFinder/, QSLiMFinder/ or SLiMProb and the default RunID is the date and time of the run.

TarGZ and SaveSpace

Obviously, the results directory can quickly fill up with files if there are multiple datasets and/or runs with different RunIDs. The way to get round this is to use the targz=T and savespace=X options.

targz=T will package up all of the files associated with a specific run into a single Dataset.RunID.tgz file. This does not work on Windows. (Note that previous versions generated a Dataset.tar.gz file.) The *.pickle.gz file associated with the run will not be included in the tar file unless savespace=2+ (see below).

Note: the tar file is actually generated from the run directory, not the results directory and will include the relative path to ResDir in the tarred files. This means that if you enter ResDir/ and then tar -xzf Dataset.RunID.tgz, an additional ResDir/ will be created in which the files can be found. This is actually pretty useful as it allows the user to unpack individual runs and then delete the whole directory when finished. To return individual results to their “rightful” place, simply run the tar command from the same directory that the SLiMSuite program was run from (e.g. tar -xzf ResDir/Dataset.RunID.tgz).

The savespace=X option saves space by deleting excess files. It is strongly recommended that this is used in conjunction with the targz=T. There are now four levels of savespace=X:

  • 0 = Delete no files
  • 1 = Delete all bar *.upc and *.pickle (Pickle excluded from tar.gz with this setting)
  • 2 = Delete all bar *.upc files (Pickle included in tar.gz with this setting)
  • 3 = Delete all dataset-specific files including *.upc and *.pickle (not *.tar.gz)

Another way to think of this is that 0 will delete nothing, 1 will leave enough files to rerun the same dataset/SLiMBuild combination, 2 will leave enough to run the same dataset with additional SLiMBuild settings, whilst 3 will cleanup absolutely everything.

The recommended setting for running on a cluster or supercomputer is targz=T savespace=1 unless file numbers are an issue, in which case targz=T savespace=2 would be better. targz=T savespace=3 is only really recommended when you are confident that all datasets will run to completion without issues. If there is a chance of nodes going down or walltimes being reached, it is better to keep the pickle files accessible for re-runs.

Saturday, 13 July 2013

SLiMSuite at the OMICS Group 3rd International Conference on Proteomics & Bioinformatics

If anyone is attending the OMICS Group 3rd International Conference on Proteomics & Bioinformatics this week then be sure to say hello. I am speaking on the last day in the “Computational Biology” track.. (Never the best time to talk at a conference as there is limited time for follow up but at least it is before lunch!)

SLiM Pickings: mining structural and sequence data for the prediction of short linear protein interaction motifs

Short Linear Motifs (SLiMs) are short functional protein sequences that act as ligands to mediate transient protein-protein interactions (PPI) in critical biological pathways and signaling networks. SLiMs are short (3-15aa), generally tolerate considerable sequence variation and typically have fewer than five residues critical for function. These features result in a degree of evolutionary plasticity not seen in domains and SLiMs often add new functions to proteins by convergent evolution. They also present a challenge for computational identification, making it difficult to differentiate biological signal from stochastic patterns. Despite this, discovering new SLiMs is of great interest due to their potential as therapeutic targets.

In recent years, we have made great progress in SLiM discovery, particularly through development of the SLiMSuite package of bioinformatics tools. SLiMs generally occur in structurally disordered regions of proteins and exhibit evolutionary conservation relative to other disordered residues. SLiMFinder uses this knowledge and exploits patterns of convergent evolution to predict novel, over-represented motifs within a statistical framework with high specificity. Applying this approach to a comprehensive set of human PPI data has highlighted interactome complexity and quality as the next challenges for SLiM prediction. Our latest development, QSLiMFinder (“Query” SLiMFinder) tackles some of these issues by incorporating specific interaction data to restrict the motif search space, which improves both the sensitivity and biological relevance of predictions. We are now using QSLiMFinder to combine structurally defined domain-motif interactions with large-scale PPI data to perform large-scale de novo SLiM prediction.

Monday, 29 April 2013

Second QSLiMFinder poster now on F1000 Posters

The second QSLiMFinder poster from the recent Cold Spring Harbor Laboratory "Systems Biology: Networks" meeting is now available at F1000 Posters:
  • Edwards RJ & Palopoli N. Computational prediction of short linear motifs mediating host-pathogen protein-protein interactions.
  • (I'm not sure why the last post about the other poster disappeared for a few days but it's back now!

    Thursday, 18 April 2013

    Latest QSLiMFinder poster now on F1000 Posters

    One of the QSLiMFinder posters from the recent Cold Spring Harbor Laboratory "Systems Biology: Networks" meeting is now available at F1000 Posters:
  • Palopoli N & Edwards RJ. Improved computational prediction of Short Linear Motifs using specific protein-protein interaction data.
  • With any luck, the other one will appear soon.

    Thursday, 28 March 2013

    QSLiMFinder at Cold Spring Habor Laboratory "Systems Biology: Networks" 2013

    This month saw another successful "Systems Biology: Networks" meeting held at Cold Spring Habor Laboratory, New York. SLiMSuite was well represented with two posters, which you can now view online if you like:

    1. Palopoli N & Edwards RJ. Improved computational prediction of Short Linear Motifs using specific protein-protein interaction data.
    Short Linear Motifs (SLiMs) are short segments of proteins that mediate numerous domain-motif interactions (DMI). In spite of the crucial role that they play in many biological pathways, their features and diversity remain understudied. The limited size and degenerate nature of SLiMs hinder their identification by pure de novo prediction methods, which must deal with a very large motif search space entirely determined by the parameters used to build the motifs.

    The most successful methods are built on an explicit model of convergent evolution for detecting over-represented motifs in unrelated proteins that share a common attribute. We have previously presented SLiMFinder[1] which accounts for the motif search space to statistically model the probability of observing a given prediction by chance. SLiMFinder greatly benefits from the incorporation of prior knowledge that reduces the sequence search space and increases sensitivity.

    More recently we have extended the standard algorithm to develop QSLiMFinder, a query-focused method of SLiM discovery. In QSLiMFinder the search space is not built from the whole set of proteins but rather from one specific query protein or region thereof. By only looking at all putative motifs in the query that may be shared by the rest, the motif space is significantly reduced and the sensitivity is increased. Moreover, DMI data can be used to focus on a specific query region rather than in the complete protein. A major plus of QSLiMFinder is its ability to incorporate this information from three-dimensional structures of interacting proteins, like those in the database of 3D Interaction Domains (3DID)[2] or as predicted from structural data[3].

    A thorough comparative benchmark of the SLiMFinder and QSLiMFinder performances on datasets of known motifs has confirmed that the latter typically returns motifs with higher significance and produces more results that are enriched against expectation. As expected, QSLiMFinder improves sensitivity by ‘zooming-in’ in the region of interest and paves the way to mine interaction data for novel SLiMs.
    1. Edwards RJ, Davey NE, Shields DC. (2007) SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS One; 2(10):e967.
    2. Stein A, Ceol A, Aloy P. (2011) 3did: identification and classification of domain-based interactions of known three-dimensional structure. Nucleic Acids Res; 39:D718-723.
    3. Stein A, Aloy P. (2010) Novel peptide-mediated interactions derived from high-resolution 3-dimensional structures. PLoS Comput Biol. 6(5):e1000789.

    2. Edwards RJ & Palopoli N. Computational prediction of short linear motifs mediating host-pathogen protein-protein interactions.
    Short Linear Motifs (SLiMs) are short functional protein sequences that act as ligands to mediate transient protein-protein interactions (PPI) in critical biological pathways and signaling networks. SLiMs are short (3-15aa), generally tolerate considerable sequence variation and typically have fewer than five residues critical for function. These features result in a degree of evolutionary plasticity not seen in domains and SLiMs often add new functions to proteins by convergent evolution. This is particularly prevalent in viruses, which often exploit SLiMs to manipulate the molecular machinery of host cells[1].

    In recent years, the numbers of tools and algorithms for SLiM discovery has increased dramatically. Of these, SLiMFinder[2], which exploits a statistical model of convergent evolution to predict novel over-represented motifs with high specificity, repeatedly performs well in comparative studies. The size and degeneracy of SLiMs presents a challenge for computational identification, making it difficult to differentiate biological signal from stochastic patterns. SLiMs generally occur in structurally disordered regions of proteins and exhibit evolutionary conservation relative to other disordered residues, which can be exploited by SLiMFinder to reduce the sequence search space and improve predictions. We have recently developed QSLiMFinder (“Query SLiMFinder”), an extended version of the algorithm that can incorporate specific interaction data to restrict the motif search space and improve both the sensitivity and biological relevance of predictions. Whereas SLiMFinder can ask the general question of which motifs are enriched in a set of proteins that interact with a common partner[3], QSLiMFinder can specifically ask which of the motifs present in a viral protein are enriched in the set of host proteins that interact with the same host partner. By applying this to combined interactomes of host-host and host-pathogen PPI, it should be possible to identify novel candidates for viral mimicry of host SLiMs.

    1. Davey NE, Travé G, Gibson TJ (2011) How viruses hijack cell regulation. Trends Biochem. Sci. 36 (3): 159–69.
    2. Edwards RJ, Davey NE, Shields DC. (2007) SLiMFinder: a probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS One; 2(10):e967.
    3. Edwards RJ, Davey NE, O'Brien K & Shields DC (2012): Interactome-wide prediction of short, disordered protein interaction motifs in humans. Molecular Biosystems 8: 282-95.

    Thursday, 20 December 2012

    New SLiMSuite, SeqSuite and RJESuite downloads available

    Just in time for Christmas, new releases of all the downloads are available at the Edwards Lab software page. Documentation is still lagging behind but will hopefully catch up (along with a bit of an overhaul of this blog). Questions welcome in the meantime.

    In addition to QSLiMFinder 1.4, the biggest change this release is probably the upgrade of GOPHER. Version 3.x features improved organisation of output files for queries from different species in addition to a capacity to have several different multiple alignment programs run on the same orthologue sets. See the website for more info.

    Updates since last release:

    • gopher: Created.
    → Version 3.0: See archived GOPHER 1.9 and gopher_V2 2.9 for history and obselete options.
    → Version 3.0: Added organise=T/F and gopherdir=PATH for improved file organisation. Tightened savespace.
    → Version 3.0: Added compfilter=T/F for improved complexity filter and composition statistics control for *initial* BLAST.
    → Version 3.0: Changed default tree extension to *.nwk for compatibility with MEGA. Deleted _phosAlign() method.
    → Version 3.0: Added orthology ID option and alignment program to customise output further.
    → Version 3.1: Added full reciprocal best hit method. (fullrbh=T/F)

    • gopher_V2: Updated from Version 2.8.
    → Version 2.9: Deleted oldStigg() method. Added simple Reciprocal Best Hit orthology prediction.

    • qslimfinder: Updated from Version 1.2.
    → Version 1.3: Updated the output for Max/Min filtering and the pickup options.
    → Version 1.4: Added additional dictionary and list to store Query dimers and SLiMs for motif space calculations.
    → Version 1.4: Added qexact=T/F option for calculating Exact Query motif space (True) or estimating from dimers (False).

    • slimfinder: Updated from Version 4.2.
    → Version 4.3: Updated the output for Max/Min filtering and the pickup options. Removed TempMaxSetting.
    → Version 4.4: Modified to work with GOPHER V3.0.

    • rje: Updated from Version 4.3.
    → Version 4.4: Added lineFromIndex(target,file,re_index='^(\S+)\s',sortunique=False,xreplace=True).

    • rje_seq: Updated from Version 3.13.
    → Version 3.14: Added CLUSTAL Omega alignment program ['clustalo']
    → Version 3.15: Added PAGAN alignment program ['pagan'] and (hopefully) fixed minor Windows fastacmd bug.

    • rje_sequence: Updated from Version 2.1.
    → Version 2.2: Added more yeast species.

    • rje_slimcalc: Updated from Version 0.4.
    → Version 0.5: Altered to use GOPHER V3 and handle nested alignment directories.

    • rje_slimlist: Updated from Version 1.0.
    → Version 1.1: Modified to work with GOPHER V3.0 for alignments.

    Tuesday, 27 November 2012

    QSLiMFinder 1.4: quicker and more efficient - available on request

    The on-going benchmarking of QSLiMFinder has thrown up a couple of discoveries to date. The first is that, reassuringly, it appears to work. (More on this another time.) The second is that it is slow. Or, at least, it was slow.

    Thankfully, the cause of its surprisingly slow performance (compared to SLiMFinder) has been tracked down and fixed. At the same time, a (related) potential memory issue with large query sequences has also been sorted out.

    The underlying problem is unlikely to have had a large effect on the SLiM prediction itself, although this is currently under investigation. The last release of SLiMSuite was only last week and, as QSLiMFinder is not officially published and released yet, I will not be compiling a new download immediately to take advantage of the improvements. The revised code is available on request if anyone is using QSLiMFinder.

    Friday, 23 November 2012

    New SLiMSuite, SeqSuite and RJESuite releases are now available

    New releases of SLiMSuite, SeqSuite and RJESuite are now available from the Edwards Lab software page.

    Please note that the documentation (particularly the manuals) are still lagging a bit behind, so do report anything that does not make sense. The default settings also need to be verified as there is a chance that some of these may have inadvertently changed over the years. (The same core code is now used for the webservers, which often have different defaults.) Checking these along with updating and checking the servers themselves are ongoing priorities.

    A full list of updated modules is given below. As well as SLiMMaker now handling end of sequence characters, the biggest changes this release are updates to CompariMotif to (3.7) output unmatched input motifs and (3.8) improve handling of partially overlapping ambiguous positions (e.g. [AGS] and [ST]). The motivation behind both these changes is the ongoing benchmarking (and preparation for publication) of QSLiMFinder and the creation of SLiMBench for benchmarking motif prediction methods. A QSLiMFinder section has been added to the SLiMFinder Manual (section 5.4). SLiMBench is still a work in progress and will be documented in a later release.

    Updates since last release:

    • comparimotif_V3: Updated from Version 3.6.
    → Version 3.7: Added coreIC and output of unmatched motifs.
    → Version 3.8: Added overlaps=T/F : Whether to include overlapping ambiguities (e.g. [KR] vs [HK]) as match [True]
    → Version 3.8: Changed scoring of overlapping ambiguities - uses IC of all possible ambiguities. Added "Ugly" match type.

    • slimbench: Created.
    → Version 0.0: Initial Compilation.
    → Version 0.1: Functional version with benchmarking dataset generation.
    → Version 1.0: Consolidation of "working" version with additional basic benchmarking analysis.
    → Version 1.1: Added simulated dataset construction and benchmarking.
    → Version 1.2: Added MinIC filtering to benchmark assessment. Sorted beginning/end of line for reduced ELMs.
    → Version 1.3: Made SimCount a list rather than Integer. Sorted CompariMotif assessment issue.
    → Version 1.4: Added ICCut and SLiMLenCut as lists and output columns.
    → Version 1.5: Added Summary Results output table. Removed PropRes.

    • slimmaker: Updated from Version 1.0.
    → Version 1.1: Modified to work with end of line characters.

    • slimsearch: Updated from Version 1.5.
    → Version 1.6: Minor tweaks to Log output. Add option for UPC number in occ output.

    • rje: Updated from Version 4.1.
    → Version 4.2: Modified INI reading across the board to look in ../settings/ and look for defaults.ini as well as rje.ini.
    → Version 4.2: Enabled handing on -ini FILE in addition to ini=FILE.
    → Version 4.3: Added ilist and nlist types to cmdRead for objects. (Lists of integers and floats). Add ratio() function.

    • rje_blast: Updated from Version 1.13.
    → Version 1.14: Added blast.checkProg(qtype,stype) to check whether blastp setting matches sequence formats.

    • rje_db: Created.
    → Version 0.0: Initial Compilation.
    → Version 0.1: Added merge tables option.
    → Version 0.2: Miscellaneous updates to various methods.
    → Version 0.3: Minor doc tweaks and added keepFields().

    • rje_seq: Updated from Version 3.12.
    → Version 3.13: Updated sequence type checking for use with GABLAM 2.10.

    • rje_seqlist: Created.
    → Version 0.0: Initial Compilation. Based on rje_seq 3.10.
    → Version 0.1: Added basic species filtering and sequence output.
    → Version 0.2: Added upper case filtering.
    → Version 0.3: Added accnum filtering and sequence renaming.
    → Version 0.4: Added sequence redundancy filtering.
    → Version 0.5: Added newgene=X for sequence renaming (newgene_spcode__newaccXXX). NewAcc no longer fixed Upper Case.
    → Version 1.0: Upgraded to "ready" Version 1.0. Added concatenate=T and split=X options for sequence concatenation.
    → Version 1.0: Added reading of sequence type from rje_seq.py and mixed=T/F.
    → Version 1.1: Added shortName() and modified SeqDict.

    • rje_sequence: Updated from Version 2.0.
    → Version 2.1: Added re_unirefprot = re.compile('^([A-Za-z0-9\-]+)\s+([A-Za-z0-9]+)_([A-Za-z0-9]+)\s+')

    • rje_slim: Updated from Version 1.5.
    → Version 1.6: Fixed splitting bug introduced by lower case motifs.

    • rje_slimcore: Updated from Version 1.8.
    → Version 1.9: Minor modifications to Log output. Updated motifSeq() function to output unmasked sequences.

    • rje_slimlist: Updated from Version 0.6.
    → Version 1.0: Functional module with lower case motif splitting fixed and ? -> .{0,1} replacement.

    • rje_zen: Updated from Version 1.0.
    → Version 1.1: Added a few more words here and there.

    Tuesday, 15 May 2012

    Bioinformatics Postdoc Position available!

    A two-year BBSRC-funded postdoc position is now available to work in the Edwards lab developing and applying QSLiMFinder. Informal enquiries are encouraged. You can apply or get further details here. The blurb:
    You are invited to apply for the post of Research Fellow to work closely with Dr Richard Edwards on a BBSRC-funded project to develop and apply computational tools for the prediction of protein motifs that mediate protein-protein interactions.

    Many protein-protein interactions are mediated by Short Linear Motifs (SLiMs): short stretches of proteins (5-15 amino acids long), of which only a few positions are critical to function. These motifs are vital for biological processes of fundamental importance, such as signalling pathways and targeting proteins to the correct part of a cell.

    This position represents an exciting opportunity to join one of the early pioneers in the growing field of SLiM prediction. The primary objective of this project is to integrate a number of leading computational techniques to predict novel SLiMs and, in so doing, add crucial detail to protein-protein interaction networks. This will generate a valuable resource of potential SLiMs, including defined occurrences and interactions.

    The project will use a number of computational and sequence analysis techniques. Basic programming skills are essential. Experience with database design, HPC and web programming are desirable. You will be required to develop a thorough knowledge of SLiM-mediated protein-protein interactions and should therefore be comfortable with biological literature, biochemistry, molecular evolution and structural biology.

    A background in either computer science or biology, with a PhD in a relevant subject area, is essential. Previous research experience (PhD or Postdoctoral) in computational biology is highly desirable. Candidates with a computer science background must demonstrate an interest and aptitude for molecular biology. Similarly, candidates with a biology background must demonstrate an interest and aptitude for computer programming.

    You should be an enthusiastic researcher, a good team-worker and an excellent communicator. Project management skills and independent research experience are desirable.

    The position is full-time and available immediately for a period of up to two years.

    The closing date for this position is 15 June 2012. Please apply online through www.jobs.soton.ac.uk or alternatively telephone 023 8059 2750 for an application form. Please quote reference number 119512BJ on all correspondence. In addition to submitting your CV, please enclose a personal statement highlighting your research interests and experience, as outlined in the accompanying Further Particulars. Please note that the project is 100% computational.

    Wednesday, 9 May 2012

    SLiMSuite servers and programs

    An emerging field of biology is the role of intrinsically disordered regions in protein function and, specifically, protein-protein interactions (PPI) [1-2]. Of particular interest, Short, Linear Motifs (SLiMs) playing a vital role in disorder-mediated PPI, acting as ligands for molecular signalling, post-translational modifications and subcellular targeting [3]. SLiMs have extremely compact protein interaction interfaces, generally encoded by less than 4 major affinity-/specificity-determining residues within a stretch of 2-10 residues [4]. Their small size enables high functional density and evolutionary plasticity, which is frequently exploited by rapidly evolving pathogens that use them to hijack cellular processes [5]. These same features also make experimental discovery a challenge and considerable attention has therefore been given to computational methods for SLiM prediction and analysis [6].

    A number of these tools have been developed by the Edwards and Shields labs [7-11] and made available as part of the SLiMSuite package and online as webservers (http://bioware.ucd.ie) [9-10,12-14], with two new tools, SLiMPrints and QSLiMFinder, currently in preparation for submission, and SLiMMaker to be added soon. The main tools that form the SLiMSuite package/servers are as follows:
    • SLiMFinder [8,13]: de novo SLiM prediction based on a statistical model of over-represented motifs in unrelated proteins.
    • SLiMDisc [7,12]: de novo SLiM prediction based on heuristic ranking of over-represented motifs in unrelated proteins.
    • SLiMPred [11]: de novo SLiM/MoRF prediction in single proteins based machine learning of motif attributes.
    • SLiMSearch [10]: biological context (disorder & conservation) for searches of pre-defined motifs with under- and over-representation statistics, correcting for evolutionary relationships.
    • SLiMSearch 2.0 [14]: biological context (disorder & conservation) and ranking for proteome-wide searches of pre-defined motifs.
    • SLiMPrints (in prep.): de novo SLiM/MoRF prediction in single proteins from statistical clustering of conserved disordered residues.
    • QSLiMFinder (server coming soon): Query-based variant of SLiMFinder with increased sensitivity and specificity.
    • CompariMotif [9]: Motif-motif comparison tool.
    • SLiMMaker (coming soon): Simple tool for converting aligned peptides or SLiM occurrences into a regular expression motif.
    • GOPHER [12]: Automated orthologue prediction and alignment algorithm. Used for conservation-based masking (SLiMFinder/SLiMSearch) and prediction (SLiMPrints).
    • GABLAM [7] (server coming soon): BLAST-based protein similarity scoring and clustering. Used for SLiMFinder and SLiMSearch adjustments for evolutionary relationships.
    Personnel (and funding applications) permitting, a number of improvements for these resources are planned, including updates to the underlying databases for proteome-wide predictions (SLiMSearch 1.0 & 2.0), conservation analyses (SLiMSearch 1.0 & 2.0, SLiMPrints, GOPHER) and SLiM comparisons (CompariMotif). We also intend to improve the integration of different tools, allowing seamless continuation of analyses. Motif predictions ((Q)SLiMFinder/SLiMPrints/SLiMPred) will be able to be searched directly against known motifs (CompariMotif) or proteomes (SLiMSearch); GOPHER alignments will be accessible for SLiMPrints analyses and even SLiMSearch/(Q)SLiMFinder input; outputs of motif occurrences ((Q)SLiMFinder/SLiMSearch) can be used to redefine motifs using SLiMMaker etc. If you have any other suggestions for improvements, please let us know.


    References:
    [1] Tompa P (2011) Unstructural biology coming of age. Curr Opin Struct Biol 21: 419; [2] Babu MM et al. (2011) Intrinsically disordered proteins: regulation and disease. Curr Opin Struct Biol 21:432; [3] Diella F et al. (2008) Understanding eukaryotic linear motifs and their role in cell signaling and regulation. Front Biosci 13:6580; [4] Davey NE et al. (2012) Attributes of short linear motifs. Mol Biosyst 8:268; [5] Davey NE, Trave G & Gibson TJ (2011) How viruses hijack cell regulation. Trends Biochem Sci 36:159; [6] Davey NE, Edwards RJ & Shields DC (2010) Computational identification and analysis of protein short linear motifs. Front Biosci 15:801; [7] Davey NE, Shields DC & Edwards RJ (2006): SLiMDisc: short, linear motif discovery, correcting for common evolutionary descent. Nucleic Acids Res. 34:3546; [8] Edwards RJ, Davey NE & Shields DC (2007): SLiMFinder: A probabilistic method for identifying over-represented, convergently evolved, short linear motifs in proteins. PLoS ONE 2:e967; [9] Edwards RJ, Davey NE & Shields DC (2008): CompariMotif: Quick and easy comparisons of sequence motifs. Bioinformatics 24:1307; [10] Davey NE et al. (2010): SLiMSearch: a webserver for finding novel occurrences of short linear motifs in proteins, incorporating sequence context. Lecture Notes in Bioinformatics 6282:50; [11] Mooney C et al. (2012): Prediction of short linear protein binding regions. J Mol Biol 415:193; [12] Davey NE, Edwards RJ & Shields DC (2007): The SLiMDisc server: short, linear motif discovery in proteins. Nuc Acids Res 35:W455; [13] Davey NE et al. (2010): SLiMFinder: a web server to find novel, significantly over-represented, short protein motifs. Nuc Acids Res 38:W534; [14] Davey NE et al. (2011): SLiMSearch 2.0: biological context for short linear motifs in proteins. Nuc Acids Res 39:W56.