# What is the return?

**URL:** <https://datachallenge.cfm.fr/t/what-is-the-return/42>\
**Category:** Data\
**Created:** [February 5, 2018, 7:16pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42 "2018-02-05T19:16:32Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Twice22](https://avatars.discourse-cdn.com/v4/letter/t/e36b37/32.png) [@Twice22](https://datachallenge.cfm.fr/u/Twice22)\
**Post date:** [February 5, 2018, 7:16pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/1 "2018-02-05T19:16:32Z")

</div>

Hello.

I’m trying to understand the data of the challenge. Actually I still don’t understand what is the return. In the [video](https://www.college-de-france.fr/site/stephane-mallat/Prediction-de-volatilite-de-marches-financiers-par-CFM.htm) at 3:05 it is said that it indicates wether or not the price increased during the next 5 minutes.  
Well actually it could have make sense but when plotting the graph with the data, here is what I get:  
 ![return](https://cdck-file-uploads-canada1.s3.dualstack.ca-central-1.amazonaws.com/flex029/uploads/cfm/original/1X/8b0340cb0b750290b297f404b9eaeb118a2eebfc.png)

where:

- return = -1 --\> red dot
- return = 1 --\> green dot
- return = 0 --\> blue dot

So actually you can see red dots while the price actually increases over the **next** 5 minutes and green dots while the price actually decreases over the last 5 minutes. In the same manner if we assume that the return actually indicates the direction of the price over the **last** 5 minutes… it doesn’t work…

So, I don’t understand… **What is the meaning of the return**?

**Note** : for reproductibility purpose, here is a piece of code that will plot the same graph:

```python
import pandas as pd
df = pd.read_csv('data/training_input.csv', sep=';')

X = np.array(df.iloc[0, 3:57])
classes = np.array(df.iloc[0, 57:111])
plt.figure(figsize=(15,10))
plt.plot(X)
plt.scatter(np.where(classes == 1)[0], X[classes==1], c='g', s=50)
plt.scatter(np.where(classes == -1)[0], X[classes==-1], c='r', s=50)
plt.scatter(np.where(classes == 0)[0], X[classes==0], c='b', s=50)
plt.show()

```

---

<div class="post-metadata">

**Author:** ![AiOpH](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/aioph/32/13_2.png) [@AiOpH](https://datachallenge.cfm.fr/u/AiOpH)\
**Post date:** [February 6, 2018, 10:44am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/2 "2018-02-06T10:44:11Z")

</div>

Hello Twice22,  
The data of the training set you are plotting are **volatilities** (df.iloc[0, 3:57]), not prices.  
I think that return do indicates if prices increase or decrease but is not directly related to volatilities increase or decrease. Volatility is measure of the “spread” of the prices. The fluctuations of the price is not directly related to the value of the price. For instance, you can have a decrease in volatility but still an increase of prices thus positive return.  
If you look carefully at the graph in the video, it shows prices, not volatilities.  
So at the end, you reasoning is right for prices but not for volatilities.  
AiOpH.

---

<div class="post-metadata">

**Author:** ![Twice22](https://avatars.discourse-cdn.com/v4/letter/t/e36b37/32.png) [@Twice22](https://datachallenge.cfm.fr/u/Twice22)\
**Post date:** [February 6, 2018, 11:07am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/3 "2018-02-06T11:07:15Z")

</div>

Ah yes thank you. My bad I made a confusion between volatilities and prices

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [February 6, 2018, 11:07am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/4 "2018-02-06T11:07:56Z")

</div>

Thank you @AiOpH, this is a good description of the situation.

The key point is that a high volatility during some time interval means large price variations in this interval. This says nothing about the price change.

However minute the price change is, it has a sign, which is what is reported in the “return” columns.

@Twice22: `df.columns[3:57]` will confirm that your curve displays volatilities, and therefore only price variation amplitudes (not returns).

---

<div class="post-metadata">

**Author:** ![chaitanya](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chaitanya](https://datachallenge.cfm.fr/u/chaitanya)\
**Post date:** [September 28, 2019, 4:45pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/5 "2019-09-28T16:45:44Z")

</div>

How to Visualize this Data?

---

<div class="post-metadata">

**Author:** ![chaitanya](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chaitanya](https://datachallenge.cfm.fr/u/chaitanya)\
**Post date:** [September 28, 2019, 4:46pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/6 "2019-09-28T16:46:22Z")

</div>

How to Visualize the Data?

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [September 30, 2019, 10:46am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/7 "2019-09-30T10:46:51Z")

</div>

You can run the code in the original post in a Jupyter notebook. The only missing part is that you need to use Python and also to import the plotting module, which can be done with

```
from matplotlib import pyplot as plt
```

---

<div class="post-metadata">

**Author:** ![chaitanya](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chaitanya](https://datachallenge.cfm.fr/u/chaitanya)\
**Post date:** [October 28, 2019, 2:52pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/8 "2019-10-28T14:52:54Z")

</div>

Thank you for your reply…  
I have used a lgbm classification model and the metric is about 0.5204 using light GBM can you suggest any other models for greater accuracy?

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [October 29, 2019, 10:14am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/9 "2019-10-29T10:14:20Z")

</div>

I must say that this is precisely one of the questions that is implicitly asked to participants, so I will let other participants contribute to answers if they want to!

---

<div class="post-metadata">

**Author:** ![chaitanya](https://avatars.discourse-cdn.com/v4/letter/c/848f3c/32.png) [@chaitanya](https://datachallenge.cfm.fr/u/chaitanya)\
**Post date:** [November 2, 2019, 4:17am UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/10 "2019-11-02T04:17:10Z")

</div>

Can you tell me if any preprocessing is required on the stock data if yes then which columns should be preprocessed ???

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [November 4, 2019, 4:58pm UTC](https://datachallenge.cfm.fr/t/what-is-the-return/42/11 "2019-11-04T16:58:50Z")

</div>

This is yet another question that is asked from data scientists taking part in a data challenge. There are plenty of resources on the web on how to make predictions from data, in particular with machine learning: this is the best place for information.

Please use this forum only for questions that cannot be answered by searching the web. Thanks!
