Zobrazují se příspěvky se štítkemhandson. Zobrazit všechny příspěvky
Zobrazují se příspěvky se štítkemhandson. Zobrazit všechny příspěvky

pátek 6. října 2017

Chasing own tail

Photo Taro the Shiba Inu @ Flickr

It's a depressive moment in everyone's life, chasing own tail. What this about in my case? When I want to move in career towards data science role. Not a big data, i.e., working with Hadoop, but real data science which is about statistics, data mining, data analysis, looking for patterns, feature extraction and so on. The issue is the  experience on some project.

Yes you most likely know it's this pattern:

  1. There is no job without project experience
  2. You can't get project experience without job
Oh, wait maybe I can break this infinite loop. How?

neděle 19. března 2017

Continuous and discrete time series visualized together

The second publication for Gigascience is nearly finished and the deadline for third publication is approaching really quickly so I needed to start with data analysis. Because I had just a few ideas about how to analyze experiment data I asked the colleague from University about advice and sent her data. It was a few months ago.

Recently I got back from her the ideas and notes and also the recommendation for the book about time series analysis. As usual, it's good to start with several visualizations and from simple things to complex. So, I want to share some notes about it.

Data

As the output from two parts of the experiment were recorded two pairs of datasets for:
  • Experiment #1 with Fitbit Charge HR
    • 1029 tweets with average of 20.56 tweets per day, 
    • 411 799 records of HR with frequency of 6 - 7 records per minute
  • Experiment #2 with Peak Basis
    • 1017 tweets with average of 20.32 tweets per day,
    • 69 909 records of HR with frequency of 1 record per minute
All the tweets are just a records of sentiment and they are evaluated by hashtags #p for positive and #n for negative sentiment by experiment participant. As a part of publication, the sentiment will be extracted via machine learning methods and compared to human evaluation.

Visualization

I took data from experiment #1 since there are more records of HR and did following data wrangling:

  • Tweets respectively evaluated sentiment needs to be extended to not represent just a point in time, but the whole window. For this, we need to define "breaking point" between two consecutive going sentiments. It's simply in the middle. This is visualized in the following figure as the gray line with the dots on it where positive sentiment = 150 of HR bpm and negative sentiment = 0 of HR bpm (scale adjustment was done, otherwise sentiment is represented by +1 and -1). Dot's here representing true records, the line is an extrapolation to the time window.
  • Heart rate is drawn with another gray line representing rapid changes over time. I have applied to it simple moving average (SMA) method to get a much smoother line for further analysis. And also cut this SMA line by extrapolated windows of sentiment where the red part represents negative sentiment and blue part positive sentiment.
Here is the result:
It looks really promising, but that's all so far. I have few ideas how to continue, but since this is a part of my third publication I will publish it first and then link the publication itself and also complete code on Github. Sharing this was just about the idea how to start with the visualization of continuous and discrete time series together.

pátek 14. října 2016

There are no shortcuts

During last 3 months I have realized it is really difficult to progress with something when you don't have enough knowledge and you are trying to go ahead without proper basics and good background, because you don't have time to focus on them, rather jump directly to wild water and swim. What a mistake.

I was dealing with many things simultaneously. I needed to start with second publication for my PhD study, but I worked on second experiment which needed to be included into this publication. And of course I studied two difficult and time consuming courses in the same time on Coursera and elsewhere. All of this together with big workload in work and push from my company to learn German. To much to process, to much work to do.

Problem is, nobody can control dreams you have.
There is a reason why Magical Realism has been born in Columbia.
It's a country wheres dreams and reality are conflated.
Where in the head, people fly high as Icarus.
But even Magical Realism has it's limits
and when you get to close to the sun...
your dreams may melt away.

- Narcos

I haven't fly so high as Icarus yet. I haven't fall down and let my dreams melt away yet. But I have been close. My PhD was running away from me. I haven't finish courses.

So, there are no shortcuts. Basic need was to organize and finish all of those tasks one by one with proper priority and timing. This helped me out of the bad mantra I am too busy for basics, but I cannot manage the complex things because of lack of basics.

And I returned to basics (learn particular technique or setup, install and explore new technology) at least for once a week instead of work on my PhD or watching lot a videos from MOOC courses I am focusing on small parts with hands on experience.

neděle 3. července 2016

Heart rate and sentiment experiment design

Photo GrejGuide.dk @ Flickr
After I finished first experiment and publish article in HEALTHINF 2016 conference this year in Rome I started thinking about next experiment design.

As I mentioned in previous paper improvements I tried to get a lesson from previous mistakes and improve a lot. First, steps are increasing during the day and thus they are not so much independent. Better would be to use heart rate because is totally independent. Second, I can improve my records about sentiment in timing and evaluation. And last but least important is sentiment extraction, instead of supervised learning used in previous work I would like to used unsupervised classification.

So, let's get to details.

Experiment Design

We are still looking for relation between soft data (sentiment) and hard data (measurand). In first experiment it was text recorded via twitter and footsteps. This time it's again text recorded via twitter, but instead footsteps it's heart rate which is more idenpendent.

What's are the main objectives:
  • One month experiment (30 days)
  • 20 tweets per day (600 tweets minimum)
  • Continuous heart rate measurement (24/7)
  • Effective time for heart rate measurement between 7 and 23, i. e. 16 hours a day, rest used for wristband charging
  • No sleep activity monitoring
  • Steps monitoring? Perhaps.

neděle 3. ledna 2016

Every Data Scientist must only pay taxes and die, rest is just optional...

Picture from cacm.acm.org
I am just wondering. There are many articles about what every Data Scientist MUST know and do to be real Data Scientist. There are many articles about MUST not do as Data Scientist. What real Data Scientist MUST read and so on and so on. And honestly I don't care. Why is that so?

Research

Let's take the last one must: "What Data Scientist must read" list. I just briefly took several results from Google, here is the list of 10 of them:

And what is this quick research good for? From this circa 80 books and list of several articles you get list of subjective chosen resources which lead practically to nowhere. Just several books repeating like famous Nate Silvers Signal and Noise and of course some R or other Cookbooks. So, what is conclusion?

Conclusion

This lists of books which someone else read leads me always to Vincent Granville's article Fake data science. And what you need to take from it? Pick any book you need for your field of expertise. And what should be your field of expertise? Choose some project, doesn't matter if your personal one, school or for instance from Kaggle.com. And follow up approaches which you need to for goal achievement, then pick book, course to support your path towards this goal. And by real work, life experience you will sooner or later become Data Scientist.

středa 26. listopadu 2014

Killing me softly

The title say everything. This killing combination come with project in Sweden together with previously started PhD study and my spare time study during evenings and sleepless nights, but lets take it from beginning.

At the beginning there was sustainable project with daily routine and spring. Part of the year with the most energy. So, I applied and got into PhD study. Euphoria was the mood which I had directly when I got letter with acceptation.

Then I started study and not realized how much things I need to go trough so I started with two subjects in one semester. So far so good. Lot of study materials, I can't say no. But, really interesting topics (I will blog about it sometimes later).

"Winter is coming."

- the motto of House Stark

And then I started to work on project in Sweden. And this started to be more interesting. I am not able to study on daily basis, so I am little bit behind schedule, but I am really tired from winter and dark (at 7 AM is not daylight yet and at 2 PM is no longer daylight).

So, even when I put into it big discipline and lot of effort I have more things to do. And I hope I will handle it at least partially or with some postponements. And will finalize first semester somehow reasonable.

sobota 4. října 2014

Introduction to fitness band experiment

This started as stupid idea. It was influenced by someone's project in Coursera course (Data Analysis and Statistical Inference), when I did volunteer evaluation. The other student analyzed his own data from Fitbit Flex . I almost forgot on it. Then later on in another Coursera course (Getting and Cleaning Data from Data Science Specialization) we did data processing in R and this data contained also information from accelerometer and some experiment with couple of human subjects (see Human Activity Recognition Using Smartphones Data Set).

Then continued when I was discussing my dissertation thesis topic with my supervisor and we came to the experiment with fitness bands and sentiment written into twitter.

neděle 31. srpna 2014

Data Science and PhD study

When I was applying for the PhD study I was thinking about it as opportunity to have chance to use all techniques, programming languages, methods, processes and technologies in practical way and delivery something reasonable which supports my professional growth and development. And it could stands as proof that I am really keen to learn and improve in this field and I have passion to do Data Science.

I was thinking about it again when I have read this article The Modern Data Nerd Isn’t as Nerdy as You Think on Wired about Data Nerds who are currently in charge of Data Science departments or in similar role and apply their experiences and knowledge and usually they do not have PhD. More over they do not have master degree, just bachelor degree. Or in case they have any kind of degree it is not from field close to Data Science.

pondělí 13. ledna 2014

Looking for first hands on experience

Last week I was quite busy with MongoDB course. And even though I have plenty work with final exam which contains 10 separate works I am little bit fed up of databases at all. No matter if NoSQL or SQL which I work with really intensively in last week before Christmas.

I moved my sight to the programming language and was looking for some practical hands on experience. I have been lucky and found Data Mining: Discovering and Visualizing Patterns with Python, when I was looking for new articles and resources about data science. When I found this I was so excited! It's like kill many data science skills with one stone. I can play with Python instead of the database, work with new data and visualize it. Good, let's see how it works.