# Baseline solution

**URL:** <https://datachallenge.cfm.fr/t/baseline-solution/217>\
**Category:** Modeling\
**Created:** [July 2, 2019, 8:26pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217 "2019-07-02T20:26:20Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [July 2, 2019, 8:26pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/1 "2019-07-02T20:26:20Z")

</div>

Hi everyone,

This is the baseline solution we’re trying to beat:  
[https://colab.research.google.com/drive/1OfFkPA5wDPgrOxQBVZBlKUa8UHiBMSsZ](https://colab.research.google.com/drive/1OfFkPA5wDPgrOxQBVZBlKUa8UHiBMSsZ)

It’s in Google colab so you can run it on your browser.

All the best.

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [July 3, 2019, 10:20am UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/2 "2019-07-03T10:20:27Z")

</div>

Thank you for sharing with everybody!

The accuracies in this notebook are indeed very close to those of the benchmark. Playing a bit with the hyperparameters might bost the accuracy, as could switching to very different methods (SVM, neural networks,…).

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [July 4, 2019, 8:54am UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/3 "2019-07-04T08:54:07Z")

</div>

Now, the testing in the notebook is done on dates that can be identical to dates seen in the training: the testing is therefore not done out of sample and the observed performance can be over-optimistic.

PS: I had initially mentioned stratified sampling, but what I had in mind was splitting on _dates_ when defining the test set.

---

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [July 4, 2019, 11:39am UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/4 "2019-07-04T11:39:04Z")

</div>

Thank you for the message. Unless I made a mistake, the testing in the notebook is done dropping the dates:

X\_train, X\_test, y\_train, y\_test = train\_test\_split(df\_input.drop([‘ID’, ‘eqt\_code’, ‘date’], axis=1), df\_output[‘is\_positive’], test\_size=0.2, random\_state=42)

I’ll take a look into stratified sampling to see if I can include the dates in the training.

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [July 4, 2019, 2:28pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/5 "2019-07-04T14:28:49Z")

</div>

That’s precisely why the testing set usually contains dates already seen in training. If you imagine for instance that each date had identical inputs for all stocks, then you would find the same data in both the training and testing set (since the date and stock are removed), right? This would obviously yield a very optimistic performance estimate.

---

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [July 4, 2019, 2:35pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/6 "2019-07-04T14:35:38Z")

</div>

OK, I get it. I have trusted that ‘train\_test\_split’ from scikit-learn is well done and both sets are disjoint (to be checked).

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [July 4, 2019, 2:51pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/8 "2019-07-04T14:51:53Z")

</div>

The sets are disjoint, but they can contain the same dates. So this means that the algorithms predict some stocks on a given day by knowing “in advance” what some other stocks will do on that day. Obiously, in the extreme case where all the stocks on a given day have the exact same input data, it would be easy to make predictions for that day for stocks from the test set.

A more robust testing procedure would test on dates that are _not_ in the training set.

---

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [July 4, 2019, 3:10pm UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/9 "2019-07-04T15:10:56Z")

</div>

I haven’t yet solved the problem that was raised above, but here are some results obtained by classic stacking:  
[https://colab.research.google.com/drive/1frboVbBri7JfDz4So2xUUZhMqU9m97XE](https://colab.research.google.com/drive/1frboVbBri7JfDz4So2xUUZhMqU9m97XE)

PS: I also added SVMs.

---

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [July 10, 2019, 6:32am UTC](https://datachallenge.cfm.fr/t/baseline-solution/217/10 "2019-07-10T06:32:22Z")

</div>

I have edited the link of my first post and now the training and testing sets have different dates and indeed the performance was over-optimistic (now it’s below the benchmark but not by too much).
