# Specifications and and supervised learning type

**URL:** <https://datachallenge.cfm.fr/t/specifications-and-and-supervised-learning-type/237>\
**Category:** Modeling\
**Created:** [November 17, 2019, 11:12pm UTC](https://datachallenge.cfm.fr/t/specifications-and-and-supervised-learning-type/237 "2019-11-17T23:12:22Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Salim\_Gaci](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/salim_gaci/32/48_2.png) [@Salim\_Gaci](https://datachallenge.cfm.fr/u/Salim_Gaci)\
**Post date:** [November 17, 2019, 11:12pm UTC](https://datachallenge.cfm.fr/t/specifications-and-and-supervised-learning-type/237/1 "2019-11-17T23:12:22Z")

</div>

Hello,

It might seem simple but I wonder whether we are called to use regression or classification. In fact, if we use regression, we can then deduce the class of the output (1 if it is a positive return and 0 if not, for instance). Or, we could preprocess the data, changing then the output set to 1 or 0 depending on the sign of the return, thus we would use classification algorithms. My question is : are we “authorized” to change the y\_train values or is it mandatory to give a float value ?

---

<div class="post-metadata">

**Author:** ![lebigot](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/lebigot/32/10_2.png) [@lebigot](https://datachallenge.cfm.fr/u/lebigot)\
**Post date:** [November 18, 2019, 11:41am UTC](https://datachallenge.cfm.fr/t/specifications-and-and-supervised-learning-type/237/2 "2019-11-18T11:41:02Z")

</div>

You can definitely train on anything y\_train training targets that you want (this is a general fact).

Note however that some predictions might be hard to make (because the target can be largely random), in which case predicting a value between the two classes (instead of only a positive/negative class) gives a better score.

---

<div class="post-metadata">

**Author:** ![Alonso\_Silva](https://yyz1.discourse-cdn.com/flex029/user_avatar/datachallenge.cfm.fr/alonso_silva/32/49_2.png) [@Alonso\_Silva](https://datachallenge.cfm.fr/u/Alonso_Silva)\
**Post date:** [November 19, 2019, 3:25pm UTC](https://datachallenge.cfm.fr/t/specifications-and-and-supervised-learning-type/237/3 "2019-11-19T15:25:45Z")

</div>

Yes, actually you could try to give a cost function that combines two penalization functions:  
1.- the difference between the predicted float value and the real float value and  
2.- the class of the output  
say with a parameter alpha and (1-alpha) and try to make it learn in an end-to-end fashion.
