{
  "topic": {
    "id": 21607,
    "title": "1st Place Solution Summary",
    "authorName": "idle_speculation",
    "commentCount": 33,
    "votes": 169,
    "postDate": "2016-06-11T21:57:27.570000"
  },
  "comments": [
    {
      "id": 124099,
      "authorName": "idle_speculation",
      "votes": 10,
      "postDate": "2016-06-15T14:23:13.287000",
      "content": "<p>@vtKMH, @Alpha</p>\n\n<p>(note: try refreshing the page a couple times if the math doesn't display correctly)</p>\n\n<p>Let me try to explain the gradient descent part in more detail and let's agree that a unique user location \\\\(u\\\\) is a triple (user_location_country, user_location_region, user_location_city) and a unique hotel location \\\\(h\\\\) is a pair (srch_destination_id, hotel_cluster).</p>\n\n<p>Now let's place the distinct user and hotel locations on the globe completely at random.  So each user \\\\(u_i\\\\) will have a latitude and longitude \\\\((\\phi_{i,1},\\phi_{i,2})\\\\).  Likewise, each hotel \\\\(h_j\\\\) has a latitude and longitude \\\\((\\theta_{j,1},\\theta_{j,2})\\\\).  It's possible to write a down a formula which gives the distance along the surface of a sphere:</p>\n\n<p>$$D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})=r*\\arccos(\\sin \\phi_{i,1} \\sin \\theta_{j,1} + \\cos \\phi_{i,1} \\cos \\theta_{j,1}\\sin(\\phi_{i,2}-\\theta_{j,2}) )$$ \nWhere \\\\(r\\\\) is the radius of the sphere.</p>\n\n<p>From the training set we collect all the combinations of \\\\((u_i,h_j,d_{ij})\\\\) where \\\\(d_{ij}\\\\) is the orig destination distance.  Our goal is to adjust the placement of each \\\\(u_i\\\\) and \\\\(h_j\\\\) in order to make the computed distance \\\\(D\\\\)  as close as possible to the actual distance \\\\(d\\\\).  One way to quantify this is to define a loss \\\\(L\\\\): $$L=(D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})-d_{ij})^2$$ Whatever placement of users and hotels minimizes this loss function should  make our predicted distances very close to the actuals.</p>\n\n<p>From calculus, you may recall the notion of the gradient.  The gradient of \\\\(L\\\\), denoted \\\\(\\nabla L\\\\), is just the vector formed from the partial derivatives of \\\\(L\\\\):$$\\nabla L=(\\frac{\\partial L}{\\partial \\phi_{i,1}}, \\frac{\\partial L}{\\partial \\phi_{i,2}},\\frac{\\partial L}{\\partial \\theta_{i,1}}, \\frac{\\partial L}{\\partial \\theta_{i,2}})$$</p>\n\n<p>For our purposes, the important fact is that the negative of the gradient points in the direction the loss is decreasing most rapidly.</p>\n\n<p>If the gradient isn't sounding familiar, don't panic, there is an easy geometric description of what's going on.  Since we have two points on a sphere, we can draw the geodesic great circle through them.  If predicted distance along the great circle \\\\(D\\\\) is larger than the actual distance \\\\(d\\\\), then we can think of the negative gradient as a pair of vectors emanating from the user and hotel and pointing toward each other along the shorter arc of the great circle.  Conversely, when \\\\(D\\\\) is smaller than \\\\(d\\\\) then the negative gradient is a pair of vectors pointing toward each other along the bigger arc of the great circle.</p>\n\n<p>Either way you want to think about the gradient,  taking the current position for our user-hotel pair and moving in the direction of the negative gradient should make our loss a little smaller.  All that really happens in gradient descent is that we iterate through each triple \\\\((u_i,h_j,d_{ij})\\\\) and make a very small step from our current configuration in the direction of the negative gradient.  After many, many iterations, the hope is that we end up with a configuration of users and hotels whose distances agree with the actual distances reasonably well.</p>"
    },
    {
      "id": 123455,
      "authorName": "idle_speculation",
      "votes": 10,
      "postDate": "2016-06-12T00:54:36.270000",
      "content": "<p>@aldente</p>\n\n<p>(1) Yes, the indicator is binary</p>\n\n<p>(2) Perhaps someone more familiar with the python interface can chime in, but that sounds right.  One thing to note is that you'll also want your input data sorted by group.</p>\n\n<p>(3) I built the model on a dual socket E5-2699v3 with ~700GB ram and was only able to use about half the bookings after '2014-07-01' due to memory limitations.  The training set was around 58 million rows and it took 38 hours to fit 1200 trees.  Honestly though, it was a huge waste of electricity because there was virtually no improvement over using a training set 1/10th the size.</p>"
    },
    {
      "id": 123452,
      "authorName": "idle_speculation",
      "votes": 5,
      "postDate": "2016-06-12T00:18:44.380000",
      "content": "<p>@data_js</p>\n\n<p>Yes, the features were one-hot encoded.  Even user_id with its 1,198,786 distinct values.  It might seem a little nutty to add a million columns to your data and then try fitting a linear model on it, but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>"
    },
    {
      "id": 123464,
      "authorName": "idle_speculation",
      "votes": 3,
      "postDate": "2016-06-12T02:51:28.810000",
      "content": "<p>@Laurae</p>\n\n<p>The loss distribution for gradient descent, in my experience, had a fat tail.  L-BFGS probably suffers from the same issue.</p>\n\n<p>It was possible to get much faster convergence by capping (actual-expected) term in the gradient.  Unfortunately, capping too aggressively had a negative impact on the eventual quality of the fit.  The compromise I came up with was to lower the cap very gradually, hence the long fit time.</p>\n\n<p>I'm using C++ as my gradient descent library.  The user interface leaves a lot to be desired, but the execution time is hard to beat.</p>"
    },
    {
      "id": 123444,
      "authorName": "Jun Shi",
      "votes": 4,
      "postDate": "2016-06-11T23:54:56.027000",
      "content": "<p>Thanks for sharing your approach. I have a question regarding step #2, Factorization Machines.</p>\n\n<p>Some of the variables have large cardinalities, for example, In the train set:</p>\n\n<p>user_location_region         1008, \nuser_location_city              50447, \nsrch_destination_id            59455.</p>\n\n<p>Did you one-hot encode all of them?</p>"
    },
    {
      "id": 123782,
      "authorName": "ZYE",
      "votes": 1,
      "postDate": "2016-06-14T00:45:13.407000",
      "content": "<p>[quote=vtKMH;123726]</p>\n\n<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>\n\n<p>[/quote]</p>\n\n<p>My understanding is, gradient descent is applied to iteratively solve a set of equations - here the equations are governed by the spherical law of cosines among points on a sphere. The case that you mentioned using gradient descent to minimize cost (e.g., min c(x)), is equivalently to solve a set of gradient equations (c'(x)=0) and gradient descent is serving same purpose.   </p>"
    },
    {
      "id": 123544,
      "authorName": "happycube",
      "votes": 1,
      "postDate": "2016-06-12T16:13:54.427000",
      "content": "<p>Amazing work, and thanks for sharing!</p>\n\n<hr>\n\n<p>BTW the Amazon X1 instance is over twice as big than that Xeon box.  <em>Iff</em> you need it to win a competition <em>this</em> powerfully and know it going in, it'd be worth every penny ;)  </p>\n\n<p>But usually if one has the skills, a less powerful machine is enough...</p>"
    },
    {
      "id": 123589,
      "authorName": "idle_speculation",
      "votes": 2,
      "postDate": "2016-06-12T22:59:17.617000",
      "content": "<p>@FengLi</p>\n\n<p>The choice of the &quot;rank:pairwise&quot; algorithm was motivated by the evaluation metric.  The metric belongs to a class of information retrieval metrics where learning-to-rank models such as &quot;rank:pairwise&quot; perform quite well.  Had the metric been something else, multiclass logloss for instance, I would choose a different algorithm.</p>\n\n<p>A similar approach can work for situations with thousands of classes.  It depends on whether large number of classes can be rejected up-front.  The <a href=\"https://www.kaggle.com/c/icdm-2015-drawbridge-cross-device-connections\">ICDM 2015</a> is an example with millions of potential classes where similar techniques were successful. </p>"
    },
    {
      "id": 123722,
      "authorName": "Lawrence Chernin",
      "votes": 0,
      "postDate": "2016-06-13T18:43:39.467000",
      "content": "<p>Who is the person of @idle_speculation ?</p>"
    },
    {
      "id": 514922,
      "authorName": "xiaojian",
      "votes": 0,
      "postDate": "2019-04-12T03:13:35.033000",
      "content": "<p>Hi , I saw your message from a instructional video . He gave you a very good rating.\nI don't know if you're going to reply to me , Because I know it's a lucky thing to accept the advice of a strong man.\nI see you've won a lot of first place . I'm very interested in data science .\n And I want to become a Data Scientists . Now my Learning route is probability theory , EDA  and Visualization .But I think I lack that kind of data science thinking . I don't know how to cultivate this kind of thinking, because I can't find any information around me.\nI look forward to your reply.</p>"
    },
    {
      "id": 292202,
      "authorName": "Mayank More",
      "votes": 0,
      "postDate": "2018-03-07T16:07:28.640000",
      "content": "<p>@idle_speculation :</p>\n\n<p>Since the factorization machines were built to exploit interactions between categorical variables, does that mean we should not train the model on numerical variables and instead always convert them to categorical variables as well?</p>"
    },
    {
      "id": 136264,
      "authorName": "TomBiernacki",
      "votes": 0,
      "postDate": "2016-09-20T21:10:46.890000",
      "content": "<p>Great job and congrats!! I am a newbie and learning so thanks for the details.</p>"
    },
    {
      "id": 132848,
      "authorName": "ec-ccs",
      "votes": 0,
      "postDate": "2016-08-29T18:32:38.777000",
      "content": "<p>Wow, @idle_speculation:</p>\n\n<p>amazing job...</p>\n\n<p>your preparation and determination to reach a high rank in this project, the fact that you took the time to explain your project to everyone after wining it, and then donating part of your prize on Lucas' honour...</p>\n\n<p>All are admirable...</p>\n\n<p>Thanks and Success!</p>"
    },
    {
      "id": 124819,
      "authorName": "Pietro Marinelli",
      "votes": 0,
      "postDate": "2016-06-22T15:17:04.403000",
      "content": "<p>Congrats and most of all thanks for sharing!!</p>"
    },
    {
      "id": 124463,
      "authorName": "Walraaf",
      "votes": 0,
      "postDate": "2016-06-18T21:30:59.147000",
      "content": "<p>Hi idle_speculation, grats on the giant win! \nYou asked for how other teams approached the distance matrix problem, so I thought I'd share what we did. We only started attempting this in the last week after I saw your score come in and I knew the suspicion I had of this being possible had to be true. We were unable to complete the approach, but what we did was attempt to identify unique hotels to improve the accuracy, by only selecting hotel cluster + destination combinations that always had consistent distances from all cities (ie two different distances from one city = more than one hotel). Then we found sets of 2 unique hotels and 2 cities that had to be on a straight line, i.e. distance a + b + c = d. This let us accurately determine  the distance between city a &amp; b (we found cities &amp; hotels that were apart by hundreds of miles while lying within a foot of a straight line). The main issue we then had was that even though we were using the WGS ellipsoid for our model of the earth instead of a sphere, we needed to know some initial coordinates, otherwise the curvature of the earth would always mess with our accuracy. I never got past this point, but looking at your elegant solution now with gradient descent, I wonder whether it would've been possible to use that to initialize the location of the first few coordinates, and triangulate everything with high accuracy from there.</p>\n\n<p>In any case, cheers for the great solution and thanks for the explanation.</p>"
    },
    {
      "id": 124301,
      "authorName": "idle_speculation",
      "votes": 0,
      "postDate": "2016-06-16T23:45:50.130000",
      "content": "<p>@Watts</p>\n\n<p>By &quot;burst&quot; I just mean that each booking instance was repeated 100 times, once for each hotel cluster.  This was done because learning-to-rank models have an additional group structure in which comparisons are made.  In this case, the group consists of a single booking instance and the items to be ranked are the hotel clusters.</p>\n\n<p>@CPMP</p>\n\n<p>You're right the distance minimization problem is very far from convex.  Even after all the parameter tuning, my final solution was definitely not the global minimum.  One technique used to mitigate this problem, which was glossed over in the summary, was to run both versions of the matrix completion 18 times in parallel with different seeds.  The average and minimum distances across the different seeds were what actually got used in the model.   </p>"
    },
    {
      "id": 124213,
      "authorName": "CPMP",
      "votes": 0,
      "postDate": "2016-06-16T09:21:07.567000",
      "content": "<p>@idle_speculation, adding kudos to all you received.</p>\n\n<p>The distance minimization problem you describe isn't convex.  Were you stuck in a local minima?</p>"
    },
    {
      "id": 124209,
      "authorName": "Ashish Lal",
      "votes": 0,
      "postDate": "2016-06-16T08:59:48.487000",
      "content": "<p>@idle_speculation</p>\n\n<p>Congratulations!</p>\n\n<p>Can u also explain how you used &quot;bursting&quot; to convert the problem into xgboost &quot;rank:pairwise&quot; problem. I am unable to understand that part.</p>\n\n<p>Thanks!</p>"
    },
    {
      "id": 123853,
      "authorName": "Beta",
      "votes": 0,
      "postDate": "2016-06-14T06:57:28.947000",
      "content": "<p>Congrats for this winning and for your helping hand .... :-) <br>\nI want to learn more about distance matrix completion part .... \nAny links to understand the theory and code of this part will be highly helpful .... \nWas your complete intention was to hot encode everything other than numeric variables and then reduce the space into lower dimentional space and then using xgboost </p>\n\n<p>Thank u very much for your help </p>"
    },
    {
      "id": 123726,
      "authorName": "vtKMH",
      "votes": 0,
      "postDate": "2016-06-13T18:52:49.890000",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>"
    },
    {
      "id": 123651,
      "authorName": "Alessandro Mariani",
      "votes": 0,
      "postDate": "2016-06-13T10:09:27.267000",
      "content": "<p>&quot;Chapeau&quot; for the solution, sharing and mostly donating! Fantastic solution :)</p>"
    },
    {
      "id": 123576,
      "authorName": "Davut Polat",
      "votes": 0,
      "postDate": "2016-06-12T20:15:01.430000",
      "content": "<p>Thanks for the details,\nthis is an epic win!  very well deserved position</p>"
    },
    {
      "id": 123555,
      "authorName": "FengLi",
      "votes": 0,
      "postDate": "2016-06-12T17:04:41.730000",
      "content": "<p>@ idle_speculation  Congratulations, thank you for sharing and I'm impressed with what you did to American Cancer Society in honor of Lucas. </p>\n\n<p>For the solution, I have one question. I noticed you that you just made only one submission. How did you know that converting this problem into <em>rank:pairwise</em> problem will help you a lot? Is there any efficient way to solve the similar problem which has thousands of classes?  Thank you.</p>"
    },
    {
      "id": 123517,
      "authorName": "idle_speculation",
      "votes": 0,
      "postDate": "2016-06-12T12:30:58.550000",
      "content": "<p>@Tagarin</p>\n\n<p>My understanding of FM models is that they are motivated by <a href=\"https://en.wikipedia.org/wiki/Non-negative_matrix_factorization\">matrix factorization</a> which one could consider as a sort of dimension reduction technique for interacting pairs of categorical variables.  </p>\n\n<p>In terms of increasing the dimensionality, it does improve results up to a point.  Past that point overfitting tends to become an issue.  The threshold in question will depend on the data.</p>"
    },
    {
      "id": 123479,
      "authorName": "Gert",
      "votes": 0,
      "postDate": "2016-06-12T06:58:27.680000",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>the difference between the matrix completion predicted distance and the distance provided for the <strong><em>j</em></strong>-th booking</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for your nice summary idle_speculation, very elegant how you included those matrix completed distances as a feature - an excellent way to deal with their imprecision. Congratulations on the extremely big win!</p>"
    },
    {
      "id": 123467,
      "authorName": "Tagarin",
      "votes": 0,
      "postDate": "2016-06-12T03:32:36.830000",
      "content": "<p>Thanks for sharing, @idle_speculation.</p>\n\n<p>[quote=idle_speculation;123452]</p>\n\n<p>but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>\n\n<p>[/quote]</p>\n\n<p>That's very interesting. This brings up two interesting questions:</p>\n\n<ul>\n<li>How would this compare with other dimensionality reduction\ntechniques? Like PCA.</li>\n<li>Would increasing dimensionality to more than 4 give better results?</li>\n</ul>\n\n<p>Congratulations with winning by a big margin!</p>"
    },
    {
      "id": 123460,
      "authorName": "gmobaz",
      "votes": 0,
      "postDate": "2016-06-12T02:13:32.383000",
      "content": "<p>@idle_speculation, impressive work and congrats!</p>"
    },
    {
      "id": 123458,
      "authorName": "Jun Shi",
      "votes": 0,
      "postDate": "2016-06-12T01:43:11.667000",
      "content": "<p>@idle_speculation. Thanks for your reply.\nYou actually encoded user_id! \nBoth your algorithm and hardware specs are impressive.</p>"
    },
    {
      "id": 123457,
      "authorName": "aldente",
      "votes": 0,
      "postDate": "2016-06-12T01:37:39.593000",
      "content": "<p>@idle_speculation, thank you so much. </p>\n\n<p>~700GB RAM, wow! How do you choose representative 1/10th the size? Randomly or stratified sampling? </p>"
    },
    {
      "id": 123447,
      "authorName": "",
      "votes": 0,
      "postDate": "2016-06-12T00:06:36.817000",
      "content": ""
    },
    {
      "id": 123446,
      "authorName": "",
      "votes": 0,
      "postDate": "2016-06-12T00:05:26.720000",
      "content": ""
    },
    {
      "id": 148119,
      "authorName": "",
      "votes": 0,
      "postDate": "2016-12-03T06:22:34.257000",
      "content": ""
    },
    {
      "id": 137034,
      "authorName": "",
      "votes": 0,
      "postDate": "2016-09-27T17:21:17.980000",
      "content": ""
    }
  ],
  "index": {
    "id": "21607",
    "title": "1st Place Solution Summary",
    "authorName": "idle_speculation",
    "commentCount": "33",
    "votes": "169",
    "postDate": "2016-06-11 21:57:27.570000"
  }
}