Tuesday, 6 November 2012

Introduction to Support Vector Machines In R


Support vector machines (SVM's) have become popular among data miners, being especially suited for extreme data intensive tasks like image classification, biosequence processing, handwriting recognition, etc., but less prone to over fitting than other machine learning methods.  Dr. Lutz Hamel, author of "Knowledge Discovery with Support Vector Machines",  presents his online course "Introduction to Support Vector Machines In R" at Statistics.com. For more details please visit at http://www.statistics.com/SVM/.

Support vector machines (SVMs) have established themselves as one of the preeminent machine learning models for classification and regression over the past decade or so, frequently outperforming artificial neural networks in task such as text mining and bioinformatics.

"Support Vector Machines in R" teaches you what is going on "under the hood" when you use SVM's.  After completing this course, you will be able to interpret the performance of SVM models, choose model parameters well during the model evaluation and selection cycle, know how linear, polynomial, and Gaussian kernels differ, and know how to tune their parameters. In addition, you will gain a deep understanding of how the cost constant "C" affects the quality of your models.

The course is based on the R statistical computing environment.  However, the knowledge gained here is easily transferred to other knowledge discovery environments.

Who Should Take This Course:
Statisticians and data miners who need to know a variety of methods for classification.

Course Program:

Course outline: The course is structured as follows

SESSION 1: The Foundations
  • What is Knowledge Discovery?
  • Describing Data Mathematically
  • Linear Decision Surfaces and Functions
  • Perceptron Learning
    • Duality
  • Maximum Margin Classifiers
    • Quadratic Programming

SESSION 2: Support Vector Machines
  • The Lagrangian Dual
  • Dual Maximum Margin Optimization
  • Linear/Non-Linear SVMs
    • "The Kernel Trick"
  • Soft-margin Classifiers

SESSION 3: Model Evaluation and Selection
  • Performance metrics
    • the Confusion Matrix
  • Model Evaluation
    • Hold-out
    • Leave-one-out
    • N-fold Cross-validation
  • Confidence Intervals
  • Elements of Statistical Learning Theory
    • the VC-dimension
    • Empirical Risk Minimization
    • VC-confidence
    • Structural Risk Minimization

SESSION 4: Extensions to the Basic Model
  • Multi-class Classification
    • One-versus-the-rest Classification
    • Pairwise Classification
  • Regression with SVMs
    • Regression with Maximum Margin Machines
    • Regression with Support Vector Machines
    • Model Evaluation

Dr. Lutz Hamel teaches at the University of Rhode Island and founded the machine learning and data mining group there. Prior to his academic post, Dr. Hamel was Director of Software Development at Thinking Machine Corporation, and Vice President of R&D for Bluestreak, where he oversaw the development of advanced technologies for online ad delivery and optimization, and directed the building of a next generation data warehouse-driven system for campaign analysis and design tools.  Participants can ask questions and exchange comments with Dr. Hamel via a private discussion board throughout the course.

You will be able to ask questions and exchange comments with Dr. Lutz Hamel via a private discussion board throughout the course.   The courses take place online at statistics.com in a series of 4 weekly lessons and assignments, and require about 15 hours/week.  Participate at your own convenience; there are no set times when you must be online. You have the flexibility to work a bit every day, if that is your preference, or concentrate your work in just a couple of days.

For Indian participants statistics.com accepts registration for its courses at special prices in Indian Rupees through its partner, the Center for eLearning and Training (C-eLT), Pune.

For India Registration and pricing, please visit us at www.india.statistics.com.

For More details contact at
Call: 020 66009116

Websites:

Monday, 5 November 2012

Bayesian Regression Modeling via MCMC Techniques


In 1946, the Polish mathematician Stanislaw Ulam lay in bed with a fever, and passed the time attempting to calculate the probability that you would win at solitaire.  He gave up, but decided a better way would be to actually play the game 100 times, and find out how often you were successful.  This was his inspiration for what he termed the Monte Carlo method, after the casino where his uncle gambled.

Monte Carlo simulation methods have many success stories, among them the blossoming of Bayesian statistics via Markov Chain Monte Carlo (MCMC) computations.  Our online course "Bayesian Regression Modeling via MCMC Techniques" will be offered by Prof. Peter Congdon, a noted author in this area. For more details please visit at http://www.statistics.com/MCMC/.

Aim of the course:
Participants in "Bayesian Regression Modeling Via MCMC will learn how to apply Markov Chain Monte Carlo techniques to Bayesian statistical modeling using WINBUGS and R software. Topics covered include Gibbs sampling and the Metropolis-Hastings method.  Participants will learn how to implement linear regression (normal and t errors), poisson and loglinear regression, and binary/binomial regression using WinBUGS. 

Who Should Take This Course:
Statisticians and analysts who need to build statistical models of data.

Course Program:

Course outline: The course is structured as follows

SESSION 1: Using Markov Chain Monte Carlo

  • Monte Carlo vs MCMC
  • Estimating parameters and probabilities from complex models
  • Sampling from random variables
  • Gibbs sampling & full conditional densities
  • Convergence
  • Metropolis-Hastings method

 

SESSION 2:

  • Sampling from standard densities
  • Specifying priors and likelihoods
  • Assessing convergence
  • Estimating parameters, probabilities and other model based quantities: Case Studies
  • Posterior summaries

 

SESSION 3: Linear Regression Modeling in WinBUGS

  • Linear regression model in WinBUGS
  • Setting priors on regression coefficients and residual variances
  • Predictor selection
  • Extending the Normal linear model (outliers, heteroscedasticity)

 

SESSION 4: General Linear Modeling in WinBUGS

  • Logistic regression for binary and binomial responses; using other links
  • Poisson regression
  • Latent data approach for binary regression
  • Loglinear models for contingency tables

Dr. Peter Congdon is a Research Professor in Quantitative Geography and Health Statistics at Queen Mary University of London. He is the author of "Bayesian Statistical Modeling," "Applied Bayesian Modeling," and "Bayesian Models for Categorical Data," all published by Wiley, as well as numerous articles in peer-reviewed journals.  His research interests include spatial data analysis, Bayesian statistics, latent variable models and epidemiology.  Participants can ask questions and exchange comments with Dr. Congdon via a private discussion board throughout the period.

You will be able to ask questions and exchange comments with Dr. Peter Congdon via a private discussion board throughout the course.   The courses take place online at statistics.com in a series of 4 weekly lessons and assignments, and require about 15 hours/week.  Participate at your own convenience; there are no set times when you must be online. You have the flexibility to work a bit every day, if that is your preference, or concentrate your work in just a couple of days.

For Indian participants statistics.com accepts registration for its courses at special prices in Indian Rupees through its partner, the Center for eLearning and Training (C-eLT), Pune.

For India Registration and pricing, please visit us at www.india.statistics.com.

For More details contact at
Call: 020 66009116

Websites:

Friday, 2 November 2012

Categorical Data - Applied Modeling


Technique-driven statistical analysis can be like the blind men sent by the king to examine an elephant.  The man who feels the leg says an elephant is like a pillar, the one who feels the tails says it is like a rope, the one who feels the tusk says it is like a pipe.  It is left to the king to integrate the accounts, which he does masterfully (this tale is told by the king).

"Categorical Data - Applied Modeling," which approaches analysis from the data end - recognize the type of data you have, and then assess what different techniques are appropriate. After taking this course, students will know how to perform logistic regression (with both binomial and multinomial response), probit, logit and loglinear analysis using statistical software. Model diagnostics and interpretation of results are also covered, and longitudinal analysis is introduced. A perfect course is Dr. Brian Marx's "Categorical Data - Applied Modeling" online at Statistics.com. For more details please visit at http://www.statistics.com/categorical2/.

Who Should Take This Course:
Any researcher or analyst who encounters categorical data and needs to analyze or model it.

Course Program:

Course outline: The course is structured as follows

SESSION 1: Logistic Regression Review
  • Interpretation of parameters and odds ratio
  • Standard errors
  • Probit analysis
  • Variable selection
  • Multiple logistic regression with categorical predictors
  • Building and applying logit models
  • Model selection (backward elimination)
  • Model checking
  • Sparse data

SESSION 2: Multicategory Logit Models
  • Logistic regression with multinomial response
  • proportional odds models

SESSION 3: Loglinear Models for Contingency Tables
  • Loglinear models for independence
  • Association
  • 3-way and 4-way tables
  • Connections
  • Logit models
  • Analysis of deviance
  • Graphical modeling
  • Linear by linear association
  • Models for ordinal responses

SESSION 4: Model for Matched Pairs
  • McNemar test
  • Symmetry models
  • Quasi-symmetry
  • Ordinal quasi-symmetry
  • Quasi-independence models
  • Rater agreement
  • Bradley-Terry model

Dr. Brian Marx is Professor of Statistics at Louisiana State University, and has taught Categorical Data Analysis for over ten years. He is currently serving as Chair of the Statistical Modelling Society and is the Coordinating Editor of Statistical Modelling: An International Journal. Dr. Marx has numerous publications in peer reviewed journals.

You will be able to ask questions and exchange comments with Dr. Brian Marx via a private discussion board throughout the course.   The courses take place online at statistics.com in a series of 4 weekly lessons and assignments, and require about 15 hours/week.  Participate at your own convenience; there are no set times when you must be online. You have the flexibility to work a bit every day, if that is your preference, or concentrate your work in just a couple of days.

For Indian participants statistics.com accepts registration for its courses at special prices in Indian Rupees through its partner, the Center for eLearning and Training (C-eLT), Pune.

For India Registration and pricing, please visit us at www.india.statistics.com.

Call: 020 66009116

Websites:

Wednesday, 17 October 2012

Cluster Analysis


Identifying distinct criminal personalities, separating customers into different types, discovering the characteristics of different kinds of stars -- all involve the use of statistical clustering methods.  Learn more about the statistical underpinnings of cluster analysis and its implementation in Anthony Babinec's online course "Cluster Analysis" at Statistics.com. For more details please visit at http://www.statistics.com/clustering/.

This course will teach you how to use various cluster analysis methods to identify possible clusters in multivariate data. In marketing applications, clusters of customer records are called market segments (and the process is called market segmentation).  Clustering is also used in computational biology, environmental science, web and internet analytics, and in military and security applications.  Methods discussed include hierarchical clustering (in which smaller clusters are nested inside larger clusters), k-means clustering; two-step clustering; and normal mixture models for continuous variables.

Who Should Take This Course:
·         Marketing analysts who need to cluster customer data as part of a market segmentation strategy;
·         Computational biologists (e.g. for taxonomy);
·         Environmental scientists (e.g. for habitat studies);
·         IT specialists (e.g. in modeling web traffic patterns);
·         Military and national security analysts (e.g. in automated analysis of intercepted communications).

Course Program:

Course outline: The course is structured as follows

SESSION 1: Hierarchical Clustering

  • Hierarchical clustering - dendrograms
  • Divisive vs. agglomerative methods
  • Different linkage methods

SESSION 2: K-means Clustering


SESSION 3: Normal Mixture Model

  • Finite mixture model
  • K-means cluster as a special case

SESSION 4: Other Approaches


Before forming AB Analytics, Anthony Babinec was Director of Advanced Products Marketing at SPSS; he worked on the marketing of Clementine and introduced CHAID, neural nets and other advanced technologies to SPSS users. He has presented at the AMA's Applied Research Methods Conference and Advanced Research Techniques Forum, the Sawtooth Software Conference, Statistical Innovation's Statistical Modeling Week, and numerous professional meetings. He is on the Board of Directors of the Chicago Chapter of the American Statistical Association, where he has held various offices including President. He is on the Editorial Board of the "Journal of Targeting, Measurement and Analysis for Marketing."  Participants can ask questions and exchange comments directly with Dr. Babinec via a private discussion board throughout the period.

You will be able to ask questions and exchange comments with Mr. Anthony Babinec via a private discussion board throughout the course.   The courses take place online at statistics.com in a series of 4 weekly lessons and assignments, and require about 15 hours/week.  Participate at your own convenience; there are no set times when you must be online. You have the flexibility to work a bit every day, if that is your preference, or concentrate your work in just a couple of days.

For Indian participants statistics.com accepts registration for its courses at special prices in Indian Rupees through its partner, the Center for eLearning and Training (C-eLT), Pune.

For India Registration and pricing, please visit us at www.india.statistics.com.

Call: 020 66009116

Websites:

Monday, 15 October 2012

Avoiding Selection Bias in Randomized Clinical Trials


Do you like to learn the essential concepts required to design rigorous randomized trials so as to ensure valid treatment comparisons, primarily by avoiding selection bias and other biases. The nature and objectives of randomization are discussed, as are those of masking, allocation concealment, blocking, stratification, dynamic randomization, and various types of bias that can arise. In addition, we cover analysis techniques that can be used to salvage reliable treatment comparisons even if some of these biases are detected. These methods are more advanced, and involve adaptations of the propensity score. We round out the course with consideration of crossover designs and self-controlled studies. Learn with Dr. Vance W. Berger in his online course "Avoiding Selection Bias in Randomized Clinical Trials" at Statistics.com. For more details please visit at http://www.statistics.com/ClinBias/.

Who Should Take This Course:
Anyone who designs, conducts, analyzes, or reviews randomized clinical trials. This includes research staff in pharmaceutical and biotechnology firms and CROs, regulators, journal editors, and students of biostatistics and epidemiology.

Course Program:

Course outline: The course is structured as follows

SESSION 1: Clinical Trial Designs - A Hierarchy
  • The need for comparison groups.
  • Historical vs. parallel comparison groups.
  • Self-selection vs. physician treatment decision vs. randomization.
  • Masking and allocation concealment.

SESSION 2: Detecting Violations of the Randomization Procedure
  • Baseline comparisons by treatment groups.
  • Baseline comparisons by P groups.
  • The Berger-Exner test and graph.

SESSION 3: Preventing Subversion of the Trial by Violations of Randomization
  • Permuted blocks with fixed or varied block size.
  • The maximal procedure.
  • Unrestricted randomization and chronological bias.

SESSION 4: Salvaging Trials Affected by Violations of Randomization
  • Excluding data from contaminated centers.
  • Excluding data from patients with predictable allocations.
  • Using the RPS as a covariate.

SESSION 5: Crossover Designs
  • Washout period.

The instructor, Dr. Vance W. Berger, is an Adjunct Professor at University of Maryland Baltimore County. He has written or co-written dozens of papers in peer reviewed journals in medicine and statistics and regularly reviews manuscripts for many of the top scholarly journals. He has also taught in the past at Rutgers University and the Johns Hopkins University School of Public Health, lectured all over the country on his research, and served as an FDA reviewer for over four years.

You will be able to ask questions and exchange comments with Dr. Vance W. Berger via a private discussion board throughout the course.   The courses take place online at statistics.com in a series of 5 weekly lessons and assignments, and require about 15 hours/week.  Participate at your own convenience; there are no set times when you must be online. You have the flexibility to work a bit every day, if that is your preference, or concentrate your work in just a couple of days.

For Indian participants statistics.com accepts registration for its courses at special prices in Indian Rupees through its partner, the Center for eLearning and Training (C-eLT), Pune.

For India Registration and pricing, please visit us at www.india.statistics.com.

Call: 020 66009116

Websites: