Dimensionality Reduction on Github Event using PCA approach

This case study using Github Event dataset focus on Malaysia’s developers.

This is a read-only API to the GitHub events. These events power the various activity streams on the site.

github star wars

github star wars

The columns in this dataset are:

  1. a_login
  2. e_CommitCommentEvent
  3. e_CreateEvent
  4. e_DeleteEvent
  5. e_DeploymentEvent
  6. e_DeploymentStatusEvent
  7. e_DownloadEvent
  8. e_FollowEvent
  9. e_ForkEvent
  10. e_ForkApplyEvent
  11. e_GistEvent
  12. e_GollumEvent
  13. e_IssueCommentEvent
  14. e_IssuesEvent
  15. e_MemberEvent
  16. e_MembershipEvent
  17. e_PageBuildEvent
  18. e_PublicEvent
  19. e_PullRequestEvent
  20. e_PullRequestReviewCommentEvent
  21. e_PushEvent
  22. e_ReleaseEvent
  23. e_RepositoryEvent
  24. e_StatusEvent
  25. e_TeamAddEvent
  26. e_WatchEvent

Sample Github Event data.

sample Github Event data

sample Github Event data

Lower dimension representation of our data frame.

lower dimension representation of our data frame

lower dimension representation of our data frame

Explained variance ratio.

explained variance ratio

explained variance ratio

Plot on the data frame.

plot on the data frame

plot on the data frame

Re-scaled mean per a_login across all the events.

re-scaled mean per a_login across all the events

re-scaled mean per a_login across all the events

Bubble plot chart (a_login mean).

bubble plot chart (a_login mean)

bubble plot chart (a_login mean)

Bubble plot chart (a_login sum).

bubble plot chart (a_login sum)

bubble plot chart (a_login sum)