{
  "id": 383466,
  "title": "Data Explorer for EDA and coreset selection",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/383466",
  "author_name": "Luke Hornof",
  "post_date": "2023-02-03T20:16:40.153000",
  "votes": 20,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I work at Akridata and we recently launched <a href=\"https://akridata.ai/data-explorer\" target=\"_blank\">Data Explorer</a>, a data-centric platform for visual data. We believe great models come from great data.</p>\n<p>I ran Data Explorer on the RSNA Mammography dataset to see what it revealed. My initial results were pretty neat – images get automatically clustered based on properties like view (CC or MLO), tissue density, or whether they have an implant or not, etc.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F4b1660299106346054698e1596a81ceb%2Fkaggle1.jpg?generation=1675454585748924&amp;alt=media\" alt=\"\"></p>\n<p>You can even do things like look at all of the images with cancer. And/or compare them to the images without cancer.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F67cd26555e70b79bd5817f030dbbe05a%2Fkaggle2.png?generation=1675454627192846&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2Fb138eea3801368be2d740897017de4e0%2Fkaggle3.png?generation=1675454640048750&amp;alt=media\" alt=\"\"></p>\n<p>Data Explorer does a lot of other things, like generate precision-recall curves, view images in confusion matrices, and perform coreset selection for faster training iteration. If you’re curious about my initial findings (and how to reproduce them), check out my <a href=\"https://akridata.ai/blog/data-explorer-for-kaggle-rsna/\" target=\"_blank\">Data Explorer for Kaggle</a> blog.  Thanks!</p>",
  "messages": [
    {
      "id": 2128568,
      "postDate": "2023-02-03T20:16:40.153Z",
      "content": "<p>I work at Akridata and we recently launched <a href=\"https://akridata.ai/data-explorer\" target=\"_blank\">Data Explorer</a>, a data-centric platform for visual data. We believe great models come from great data.</p>\n<p>I ran Data Explorer on the RSNA Mammography dataset to see what it revealed. My initial results were pretty neat – images get automatically clustered based on properties like view (CC or MLO), tissue density, or whether they have an implant or not, etc.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F4b1660299106346054698e1596a81ceb%2Fkaggle1.jpg?generation=1675454585748924&amp;alt=media\" alt=\"\"></p>\n<p>You can even do things like look at all of the images with cancer. And/or compare them to the images without cancer.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F67cd26555e70b79bd5817f030dbbe05a%2Fkaggle2.png?generation=1675454627192846&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2Fb138eea3801368be2d740897017de4e0%2Fkaggle3.png?generation=1675454640048750&amp;alt=media\" alt=\"\"></p>\n<p>Data Explorer does a lot of other things, like generate precision-recall curves, view images in confusion matrices, and perform coreset selection for faster training iteration. If you’re curious about my initial findings (and how to reproduce them), check out my <a href=\"https://akridata.ai/blog/data-explorer-for-kaggle-rsna/\" target=\"_blank\">Data Explorer for Kaggle</a> blog.  Thanks!</p>",
      "rawMarkdown": "I work at Akridata and we recently launched [Data Explorer](https://akridata.ai/data-explorer), a data-centric platform for visual data. We believe great models come from great data.\n\nI ran Data Explorer on the RSNA Mammography dataset to see what it revealed. My initial results were pretty neat – images get automatically clustered based on properties like view (CC or MLO), tissue density, or whether they have an implant or not, etc.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F4b1660299106346054698e1596a81ceb%2Fkaggle1.jpg?generation=1675454585748924&alt=media =510x300)\n\nYou can even do things like look at all of the images with cancer. And/or compare them to the images without cancer.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F67cd26555e70b79bd5817f030dbbe05a%2Fkaggle2.png?generation=1675454627192846&alt=media =265x200) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2Fb138eea3801368be2d740897017de4e0%2Fkaggle3.png?generation=1675454640048750&alt=media =265x200)\n\nData Explorer does a lot of other things, like generate precision-recall curves, view images in confusion matrices, and perform coreset selection for faster training iteration. If you’re curious about my initial findings (and how to reproduce them), check out my [Data Explorer for Kaggle](https://akridata.ai/blog/data-explorer-for-kaggle-rsna/) blog.  Thanks!",
      "votes": 20
    },
    {
      "id": 2131675,
      "postDate": "2023-02-06T09:20:04.880Z",
      "content": "<p>Magic! Great tool and inspiration. I have spent a lot of time filtering data with pandas and ploting images in notebook to understand dataset.  This tool could be really very helpful. \"coreset\" option is my takeaway - learning from today 😁</p>\n<p>Do you have any standard projects (playground) where we can experiment without uploading the same data (rsna)?</p>",
      "rawMarkdown": "Magic! Great tool and inspiration. I have spent a lot of time filtering data with pandas and ploting images in notebook to understand dataset.  This tool could be really very helpful. \"coreset\" option is my takeaway - learning from today 😁\n\nDo you have any standard projects (playground) where we can experiment without uploading the same data (rsna)?",
      "votes": 3,
      "replies": [
        {
          "id": 2131701,
          "postDate": "2023-02-06T09:48:20.730Z",
          "content": "<p>Visualization and data sampling are just the beginning :)</p>\n<p>Yes - when you register, you get built-in a few datasets to play with (aside from the rsna):</p>\n<ul>\n<li>Pascal-voc12 (3000 images with 20 object classes)</li>\n<li>BDD100k-video (16 video snippets)</li>\n<li>BDD100k-images (3000 images)</li>\n</ul>",
          "rawMarkdown": "Visualization and data sampling are just the beginning :)\n\nYes - when you register, you get built-in a few datasets to play with (aside from the rsna):\n- Pascal-voc12 (3000 images with 20 object classes)\n- BDD100k-video (16 video snippets)\n- BDD100k-images (3000 images)",
          "votes": 2,
          "replies": [
            {
              "id": 2131729,
              "postDate": "2023-02-06T10:10:27.067Z",
              "content": "<p>I can see. I play with tool. Thank you!</p>",
              "rawMarkdown": "I can see. I play with tool. Thank you!",
              "votes": 1
            },
            {
              "id": 2131739,
              "postDate": "2023-02-06T10:17:09.227Z",
              "content": "<p>Happy to set a demo if you'd like - feel free to choose from here:<br>\n<a href=\"https://calendly.com/alexander-berkovich/akridata-meeting-30-min\" target=\"_blank\">https://calendly.com/alexander-berkovich/akridata-meeting-30-min</a></p>",
              "rawMarkdown": "Happy to set a demo if you'd like - feel free to choose from here:\nhttps://calendly.com/alexander-berkovich/akridata-meeting-30-min",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2132809,
      "postDate": "2023-02-07T03:51:58.240Z",
      "content": "<p>Let me add a quick note on the <a href=\"https://docs.akridata.ai/docs/overview-analyze\" target=\"_blank\">model analyze feature</a>… </p>\n<p>If you have predictions from cross validation, upload them as a CSV file to get a confusion matrix view into the data. This is a convenient way to browse true positives vs. false positives, for example.</p>\n<p><strong>True positive predictions:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2Fb772edd28e96699792ec3b4253f8da26%2Frsna-true-pos.gif?generation=1675736993267472&amp;alt=media\" alt=\"\"></p>\n<p><strong>False positive predictions:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2F1be3bdeced0d38ef70781f79961e1adc%2Frsna-false-pos.gif?generation=1675737951242086&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Let me add a quick note on the [model analyze feature](https://docs.akridata.ai/docs/overview-analyze)... \n\nIf you have predictions from cross validation, upload them as a CSV file to get a confusion matrix view into the data. This is a convenient way to browse true positives vs. false positives, for example.\n\n**True positive predictions:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2Fb772edd28e96699792ec3b4253f8da26%2Frsna-true-pos.gif?generation=1675736993267472&alt=media)\n\n**False positive predictions:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2F1be3bdeced0d38ef70781f79961e1adc%2Frsna-false-pos.gif?generation=1675737951242086&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 2157512,
          "postDate": "2023-02-24T04:40:01.480Z",
          "content": "<p>This is a really nice feature. Thanks for sharing your animations -- they demonstrate clearly how this feature works.</p>",
          "rawMarkdown": "This is a really nice feature. Thanks for sharing your animations -- they demonstrate clearly how this feature works.",
          "votes": 1
        },
        {
          "id": 2157615,
          "postDate": "2023-02-24T07:15:17.620Z",
          "content": "<p>that is a cool feature, thanks for showing that. how do you measure the [true positive predictions/false positive predictions] from the tool are correct ? by looking at each result ? </p>",
          "rawMarkdown": "that is a cool feature, thanks for showing that. how do you measure the [true positive predictions/false positive predictions] from the tool are correct ? by looking at each result ? ",
          "replies": [
            {
              "id": 2157730,
              "postDate": "2023-02-24T09:03:01.210Z",
              "content": "<p>Data Explorer allows you to analyze model training results. To achieve this, ground-truth and model output are uploaded via a csv and the confusion matrix + histogram seen above are generated.</p>\n<p>The confidence slide bar affects the conf. matrix to see results for a given conf. threshold too.</p>\n<p>Moreover, similar analysis can be done on object detection models too.</p>\n<p>Finally, to emphasize, Data Explorer doesn't ask or need access to the model, but only model output, so your code, model and architecture are secure.</p>\n<p>I invite you have a look at our website for more info:<br>\n<a href=\"https://akridata.ai/data-explorer/\" target=\"_blank\">https://akridata.ai/data-explorer/</a><br>\nand open a FREE account to work on pre-loaded or YOUR data:<br>\n<a href=\"https://subscriptions.akridata.ai/organizations/register\" target=\"_blank\">https://subscriptions.akridata.ai/organizations/register</a></p>",
              "rawMarkdown": "Data Explorer allows you to analyze model training results. To achieve this, ground-truth and model output are uploaded via a csv and the confusion matrix + histogram seen above are generated.\n\nThe confidence slide bar affects the conf. matrix to see results for a given conf. threshold too.\n\nMoreover, similar analysis can be done on object detection models too.\n\nFinally, to emphasize, Data Explorer doesn't ask or need access to the model, but only model output, so your code, model and architecture are secure.\n\nI invite you have a look at our website for more info:\nhttps://akridata.ai/data-explorer/\nand open a FREE account to work on pre-loaded or YOUR data:\nhttps://subscriptions.akridata.ai/organizations/register"
            }
          ]
        }
      ]
    },
    {
      "id": 2132682,
      "postDate": "2023-02-07T01:51:29.957Z",
      "content": "<blockquote><p>Manual inspection of data has probably the highest value-to-prestige ratio of any activity in machine learning.</p>— Greg Brockman (@gdb) <a href=\"https://twitter.com/gdb/status/1622683988736479232?ref_src=twsrc%5Etfw\">February 6, 2023</a></blockquote>\n",
      "rawMarkdown": "<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Manual inspection of data has probably the highest value-to-prestige ratio of any activity in machine learning.</p>&mdash; Greg Brockman (@gdb) <a href=\"https://twitter.com/gdb/status/1622683988736479232?ref_src=twsrc%5Etfw\">February 6, 2023</a></blockquote> <script async src=\"https://platform.twitter.com/widgets.js\" charset=\"utf-8\"></script>",
      "votes": 4,
      "replies": [
        {
          "id": 2133919,
          "postDate": "2023-02-07T16:59:39.313Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2131133,
      "postDate": "2023-02-05T22:16:21.667Z",
      "content": "<p>After the initial visualization of the dataset, you can subsample it in various ways to train only on a portion of the data.<br>\nFor example, the \"coreset\" option will preserves small clusters, but others, like the \"random\" sampling method are available too.</p>",
      "rawMarkdown": "After the initial visualization of the dataset, you can subsample it in various ways to train only on a portion of the data.\nFor example, the \"coreset\" option will preserves small clusters, but others, like the \"random\" sampling method are available too.",
      "votes": 1
    },
    {
      "id": 2128702,
      "postDate": "2023-02-04T00:28:17.137Z",
      "content": "<p>Thanks for promoting this, it is always good to find tools to enhance daily routines. Does it handle other kind of data too, or image data only?</p>",
      "rawMarkdown": "Thanks for promoting this, it is always good to find tools to enhance daily routines. Does it handle other kind of data too, or image data only?",
      "votes": 1,
      "replies": [
        {
          "id": 2128742,
          "postDate": "2023-02-04T02:14:37.713Z",
          "content": "<p>Thanks!  For now we're focused on \"visual data\", i.e. images and video. This was motivated by the importance of Computer Vision, e.g. medical imaging, autonomous vehicles.  It's also driven some of our design decisions, like the ability to display the data in the GUI, which is super useful for visual data. I hope you also find Data Explorer useful!</p>",
          "rawMarkdown": "Thanks!  For now we're focused on \"visual data\", i.e. images and video. This was motivated by the importance of Computer Vision, e.g. medical imaging, autonomous vehicles.  It's also driven some of our design decisions, like the ability to display the data in the GUI, which is super useful for visual data. I hope you also find Data Explorer useful!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2131901,
      "postDate": "2023-02-06T13:23:07.127Z",
      "content": "<p>kaggle is a platform for uploading data tool. kaggler earn their badges by creating data tool. If you can create some simple plugin that kaggler can use after they create their dataset, it will be good. kind of automatic EDA button at the data creation page. Now we have gptchat which can automatically generative data description ….</p>\n<p>i am late … chatgpt EDA<br>\n<a href=\"https://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4\" target=\"_blank\">https://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4</a></p>",
      "rawMarkdown": "kaggle is a platform for uploading data tool. kaggler earn their badges by creating data tool. If you can create some simple plugin that kaggler can use after they create their dataset, it will be good. kind of automatic EDA button at the data creation page. Now we have gptchat which can automatically generative data description ....\n\ni am late ... chatgpt EDA\nhttps://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4",
      "votes": 2,
      "replies": [
        {
          "id": 2132404,
          "postDate": "2023-02-06T19:45:59.257Z",
          "content": "<p>Indeed, an automatic EDA plugin would be ideal. That's <em>almost</em> what we've done with Data Explorer, since we've pre-ingested and featurized the Kaggle dataset for this competition. All you have to do is click on Get Started to <a href=\"https://akridata.ai/data-explorer/\" target=\"_blank\">sign-up</a> (free) and then choose \"Import public dataset &gt; Kaggle RSNA\" once you're in:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F9511bb14b2debb0513ba3f394e74cd0f%2Frsna1.png?generation=1675711956586515&amp;alt=media\" alt=\"\"></p>\n<p>This ChatGPT automatic coding stuff is super impressive. It's amazing how much it can already do, and it's only going to get better. And once AGI is achieved, it will even know better than humans what the relevant, interesting aspects of the dataset are, know the best way to visualize them, and then automatcially download the best-in-class libraries and write the code. We live in exciting times!</p>\n<p>Until then though (sadly), humans are still required to build the state of the art tools. To that end, we've loaded Data Explorer with all sorts of advanced features like <a href=\"https://docs.akridata.ai/docs/select-and-refine\" target=\"_blank\">multiple sampling options</a> (e.g. outlier, coreset, guassian), <a href=\"https://docs.akridata.ai/docs/simsearch-modes-and-controls\" target=\"_blank\">iterative similarity search</a> (including on a subset of an image), and all sorts of <a href=\"https://docs.akridata.ai/docs/overview-analyze\" target=\"_blank\">analysis features</a> (e.g. precision-recall curves, complexity matrices). You should really check it out -- I would love to hear your feedback!</p>",
          "rawMarkdown": "Indeed, an automatic EDA plugin would be ideal. That's *almost* what we've done with Data Explorer, since we've pre-ingested and featurized the Kaggle dataset for this competition. All you have to do is click on Get Started to [sign-up](https://akridata.ai/data-explorer/) (free) and then choose \"Import public dataset > Kaggle RSNA\" once you're in:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F9511bb14b2debb0513ba3f394e74cd0f%2Frsna1.png?generation=1675711956586515&alt=media =200x100)\n\nThis ChatGPT automatic coding stuff is super impressive. It's amazing how much it can already do, and it's only going to get better. And once AGI is achieved, it will even know better than humans what the relevant, interesting aspects of the dataset are, know the best way to visualize them, and then automatcially download the best-in-class libraries and write the code. We live in exciting times!\n\nUntil then though (sadly), humans are still required to build the state of the art tools. To that end, we've loaded Data Explorer with all sorts of advanced features like [multiple sampling options](https://docs.akridata.ai/docs/select-and-refine) (e.g. outlier, coreset, guassian), [iterative similarity search](https://docs.akridata.ai/docs/simsearch-modes-and-controls) (including on a subset of an image), and all sorts of [analysis features](https://docs.akridata.ai/docs/overview-analyze) (e.g. precision-recall curves, complexity matrices). You should really check it out -- I would love to hear your feedback!\n",
          "votes": 1,
          "replies": [
            {
              "id": 2132583,
              "postDate": "2023-02-06T22:09:01.827Z",
              "content": "<p><a href=\"https://www.kaggle.com/drluke\" target=\"_blank\">@drluke</a> </p>\n<p>Thanks for the link and information. I think Data Explorer is a very great and useful tool.</p>\n<p>microsoft has invested in ChatGPT. So one day soon for e.g. excel, you can type into the command bar and ask:<br>\n\"please show me the top 3 selling book\"<br>\n\"please suggest what are the books that have poor sale and what is the reason for it\"</p>\n<p>… a robot data analyst …</p>\n<p>….</p>\n<p>i can foresee a large change in data analytics industry later  </p>\n<p>….</p>\n<p>who knows one day there will be a robot kaggler that can suggest how you win competition or even win competition.</p>",
              "rawMarkdown": "@drluke \n\nThanks for the link and information. I think Data Explorer is a very great and useful tool.\n\nmicrosoft has invested in ChatGPT. So one day soon for e.g. excel, you can type into the command bar and ask:\n\"please show me the top 3 selling book\"\n\"please suggest what are the books that have poor sale and what is the reason for it\"\n\n... a robot data analyst ...\n\n....\n\ni can foresee a large change in data analytics industry later  \n\n\n....\n\nwho knows one day there will be a robot kaggler that can suggest how you win competition or even win competition.",
              "votes": 1
            },
            {
              "id": 2132590,
              "postDate": "2023-02-06T22:25:34.163Z",
              "content": "<p>Yeah, AutoML is another exciting field. Practitioners have been talking about models designing models for years, very \"meta\". But I haven't yet heard about models designing <em>datasets</em>… maybe that's next?! 🤔</p>",
              "rawMarkdown": "Yeah, AutoML is another exciting field. Practitioners have been talking about models designing models for years, very \"meta\". But I haven't yet heard about models designing *datasets*... maybe that's next?! 🤔"
            },
            {
              "id": 2132592,
              "postDate": "2023-02-06T22:32:22.337Z",
              "content": "<p>one suggestion for you.</p>\n<p>chatGPT is based on reinforcement learning of human feedback (prompt ranking).</p>\n<p>you should design you interface that you can collect use click and response in your data explorer.<br>\nthese are valuable data for training feedback system. i think deepmind will one day release customization tool based on your feedback data</p>",
              "rawMarkdown": "one suggestion for you.\n\nchatGPT is based on reinforcement learning of human feedback (prompt ranking).\n\nyou should design you interface that you can collect use click and response in your data explorer.\nthese are valuable data for training feedback system. i think deepmind will one day release customization tool based on your feedback data\n",
              "votes": 1
            },
            {
              "id": 2132595,
              "postDate": "2023-02-06T22:35:57.470Z",
              "content": "<p>\" models designing datasets\"</p>\n<p>i would think that that is labelling data as:</p>\n<ul>\n<li>useful for training, or probably noise in data</li>\n<li>category : e.g. for autonomous driving it can be like rainy weather, at night, city … for evaluation to test model weakness </li>\n</ul>",
              "rawMarkdown": "\" models designing datasets\"\n\ni would think that that is labelling data as:\n- useful for training, or probably noise in data\n- category : e.g. for autonomous driving it can be like rainy weather, at night, city ... for evaluation to test model weakness \n ",
              "votes": 1
            },
            {
              "id": 2134410,
              "postDate": "2023-02-08T01:11:26.590Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2134412,
              "postDate": "2023-02-08T01:12:16.893Z",
              "content": "<p>Data Explorer actually does this already. For example, users can create increasingly accurate searches by clicking \"thumbs up\" or \"thumbs down\" on images and iterating. Here's a great 1 minute video that shows exactly how to do this: <a href=\"https://akridata.ai/videos/powerful-way-to-search-visual-data\" target=\"_blank\">https://akridata.ai/videos/powerful-way-to-search-visual-data</a>. Check it out!</p>",
              "rawMarkdown": "Data Explorer actually does this already. For example, users can create increasingly accurate searches by clicking \"thumbs up\" or \"thumbs down\" on images and iterating. Here's a great 1 minute video that shows exactly how to do this: https://akridata.ai/videos/powerful-way-to-search-visual-data. Check it out!"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2131675,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2023-02-06T09:20:04.880000",
      "content": "<p>Magic! Great tool and inspiration. I have spent a lot of time filtering data with pandas and ploting images in notebook to understand dataset.  This tool could be really very helpful. \"coreset\" option is my takeaway - learning from today 😁</p>\n<p>Do you have any standard projects (playground) where we can experiment without uploading the same data (rsna)?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2131701,
          "author_name": "Alexander Berkovich",
          "author_url": "",
          "post_date": "2023-02-06T09:48:20.730000",
          "content": "<p>Visualization and data sampling are just the beginning :)</p>\n<p>Yes - when you register, you get built-in a few datasets to play with (aside from the rsna):</p>\n<ul>\n<li>Pascal-voc12 (3000 images with 20 object classes)</li>\n<li>BDD100k-video (16 video snippets)</li>\n<li>BDD100k-images (3000 images)</li>\n</ul>",
          "votes": 2,
          "replies": [
            {
              "id": 2131729,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-02-06T10:10:27.067000",
              "content": "<p>I can see. I play with tool. Thank you!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2131739,
              "author_name": "Alexander Berkovich",
              "author_url": "",
              "post_date": "2023-02-06T10:17:09.227000",
              "content": "<p>Happy to set a demo if you'd like - feel free to choose from here:<br>\n<a href=\"https://calendly.com/alexander-berkovich/akridata-meeting-30-min\" target=\"_blank\">https://calendly.com/alexander-berkovich/akridata-meeting-30-min</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2132809,
      "author_name": "Anil Thomas",
      "author_url": "",
      "post_date": "2023-02-07T03:51:58.240000",
      "content": "<p>Let me add a quick note on the <a href=\"https://docs.akridata.ai/docs/overview-analyze\" target=\"_blank\">model analyze feature</a>… </p>\n<p>If you have predictions from cross validation, upload them as a CSV file to get a confusion matrix view into the data. This is a convenient way to browse true positives vs. false positives, for example.</p>\n<p><strong>True positive predictions:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2Fb772edd28e96699792ec3b4253f8da26%2Frsna-true-pos.gif?generation=1675736993267472&amp;alt=media\" alt=\"\"></p>\n<p><strong>False positive predictions:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2F1be3bdeced0d38ef70781f79961e1adc%2Frsna-false-pos.gif?generation=1675737951242086&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2157512,
          "author_name": "Luke Hornof",
          "author_url": "",
          "post_date": "2023-02-24T04:40:01.480000",
          "content": "<p>This is a really nice feature. Thanks for sharing your animations -- they demonstrate clearly how this feature works.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2157615,
          "author_name": "Kefan Xu",
          "author_url": "",
          "post_date": "2023-02-24T07:15:17.620000",
          "content": "<p>that is a cool feature, thanks for showing that. how do you measure the [true positive predictions/false positive predictions] from the tool are correct ? by looking at each result ? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2157730,
              "author_name": "Alexander Berkovich",
              "author_url": "",
              "post_date": "2023-02-24T09:03:01.210000",
              "content": "<p>Data Explorer allows you to analyze model training results. To achieve this, ground-truth and model output are uploaded via a csv and the confusion matrix + histogram seen above are generated.</p>\n<p>The confidence slide bar affects the conf. matrix to see results for a given conf. threshold too.</p>\n<p>Moreover, similar analysis can be done on object detection models too.</p>\n<p>Finally, to emphasize, Data Explorer doesn't ask or need access to the model, but only model output, so your code, model and architecture are secure.</p>\n<p>I invite you have a look at our website for more info:<br>\n<a href=\"https://akridata.ai/data-explorer/\" target=\"_blank\">https://akridata.ai/data-explorer/</a><br>\nand open a FREE account to work on pre-loaded or YOUR data:<br>\n<a href=\"https://subscriptions.akridata.ai/organizations/register\" target=\"_blank\">https://subscriptions.akridata.ai/organizations/register</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2132682,
      "author_name": "Anil Thomas",
      "author_url": "",
      "post_date": "2023-02-07T01:51:29.957000",
      "content": "<blockquote><p>Manual inspection of data has probably the highest value-to-prestige ratio of any activity in machine learning.</p>— Greg Brockman (@gdb) <a href=\"https://twitter.com/gdb/status/1622683988736479232?ref_src=twsrc%5Etfw\">February 6, 2023</a></blockquote>\n",
      "votes": 4,
      "replies": [
        {
          "id": 2133919,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-02-07T16:59:39.313000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2131133,
      "author_name": "Alexander Berkovich",
      "author_url": "",
      "post_date": "2023-02-05T22:16:21.667000",
      "content": "<p>After the initial visualization of the dataset, you can subsample it in various ways to train only on a portion of the data.<br>\nFor example, the \"coreset\" option will preserves small clusters, but others, like the \"random\" sampling method are available too.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2128702,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-04T00:28:17.137000",
      "content": "<p>Thanks for promoting this, it is always good to find tools to enhance daily routines. Does it handle other kind of data too, or image data only?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2128742,
          "author_name": "Luke Hornof",
          "author_url": "",
          "post_date": "2023-02-04T02:14:37.713000",
          "content": "<p>Thanks!  For now we're focused on \"visual data\", i.e. images and video. This was motivated by the importance of Computer Vision, e.g. medical imaging, autonomous vehicles.  It's also driven some of our design decisions, like the ability to display the data in the GUI, which is super useful for visual data. I hope you also find Data Explorer useful!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2131901,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-06T13:23:07.127000",
      "content": "<p>kaggle is a platform for uploading data tool. kaggler earn their badges by creating data tool. If you can create some simple plugin that kaggler can use after they create their dataset, it will be good. kind of automatic EDA button at the data creation page. Now we have gptchat which can automatically generative data description ….</p>\n<p>i am late … chatgpt EDA<br>\n<a href=\"https://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4\" target=\"_blank\">https://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2132404,
          "author_name": "Luke Hornof",
          "author_url": "",
          "post_date": "2023-02-06T19:45:59.257000",
          "content": "<p>Indeed, an automatic EDA plugin would be ideal. That's <em>almost</em> what we've done with Data Explorer, since we've pre-ingested and featurized the Kaggle dataset for this competition. All you have to do is click on Get Started to <a href=\"https://akridata.ai/data-explorer/\" target=\"_blank\">sign-up</a> (free) and then choose \"Import public dataset &gt; Kaggle RSNA\" once you're in:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F9511bb14b2debb0513ba3f394e74cd0f%2Frsna1.png?generation=1675711956586515&amp;alt=media\" alt=\"\"></p>\n<p>This ChatGPT automatic coding stuff is super impressive. It's amazing how much it can already do, and it's only going to get better. And once AGI is achieved, it will even know better than humans what the relevant, interesting aspects of the dataset are, know the best way to visualize them, and then automatcially download the best-in-class libraries and write the code. We live in exciting times!</p>\n<p>Until then though (sadly), humans are still required to build the state of the art tools. To that end, we've loaded Data Explorer with all sorts of advanced features like <a href=\"https://docs.akridata.ai/docs/select-and-refine\" target=\"_blank\">multiple sampling options</a> (e.g. outlier, coreset, guassian), <a href=\"https://docs.akridata.ai/docs/simsearch-modes-and-controls\" target=\"_blank\">iterative similarity search</a> (including on a subset of an image), and all sorts of <a href=\"https://docs.akridata.ai/docs/overview-analyze\" target=\"_blank\">analysis features</a> (e.g. precision-recall curves, complexity matrices). You should really check it out -- I would love to hear your feedback!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2132583,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-06T22:09:01.827000",
              "content": "<p><a href=\"https://www.kaggle.com/drluke\" target=\"_blank\">@drluke</a> </p>\n<p>Thanks for the link and information. I think Data Explorer is a very great and useful tool.</p>\n<p>microsoft has invested in ChatGPT. So one day soon for e.g. excel, you can type into the command bar and ask:<br>\n\"please show me the top 3 selling book\"<br>\n\"please suggest what are the books that have poor sale and what is the reason for it\"</p>\n<p>… a robot data analyst …</p>\n<p>….</p>\n<p>i can foresee a large change in data analytics industry later  </p>\n<p>….</p>\n<p>who knows one day there will be a robot kaggler that can suggest how you win competition or even win competition.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2132590,
              "author_name": "Luke Hornof",
              "author_url": "",
              "post_date": "2023-02-06T22:25:34.163000",
              "content": "<p>Yeah, AutoML is another exciting field. Practitioners have been talking about models designing models for years, very \"meta\". But I haven't yet heard about models designing <em>datasets</em>… maybe that's next?! 🤔</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2132592,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-06T22:32:22.337000",
              "content": "<p>one suggestion for you.</p>\n<p>chatGPT is based on reinforcement learning of human feedback (prompt ranking).</p>\n<p>you should design you interface that you can collect use click and response in your data explorer.<br>\nthese are valuable data for training feedback system. i think deepmind will one day release customization tool based on your feedback data</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2132595,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-02-06T22:35:57.470000",
              "content": "<p>\" models designing datasets\"</p>\n<p>i would think that that is labelling data as:</p>\n<ul>\n<li>useful for training, or probably noise in data</li>\n<li>category : e.g. for autonomous driving it can be like rainy weather, at night, city … for evaluation to test model weakness </li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2134410,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-08T01:11:26.590000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2134412,
              "author_name": "Luke Hornof",
              "author_url": "",
              "post_date": "2023-02-08T01:12:16.893000",
              "content": "<p>Data Explorer actually does this already. For example, users can create increasingly accurate searches by clicking \"thumbs up\" or \"thumbs down\" on images and iterating. Here's a great 1 minute video that shows exactly how to do this: <a href=\"https://akridata.ai/videos/powerful-way-to-search-visual-data\" target=\"_blank\">https://akridata.ai/videos/powerful-way-to-search-visual-data</a>. Check it out!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2128568": "I work at Akridata and we recently launched [Data Explorer](https://akridata.ai/data-explorer), a data-centric platform for visual data. We believe great models come from great data.\n\nI ran Data Explorer on the RSNA Mammography dataset to see what it revealed. My initial results were pretty neat – images get automatically clustered based on properties like view (CC or MLO), tissue density, or whether they have an implant or not, etc.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F4b1660299106346054698e1596a81ceb%2Fkaggle1.jpg?generation=1675454585748924&alt=media =510x300)\n\nYou can even do things like look at all of the images with cancer. And/or compare them to the images without cancer.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2F67cd26555e70b79bd5817f030dbbe05a%2Fkaggle2.png?generation=1675454627192846&alt=media =265x200) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F482887%2Fb138eea3801368be2d740897017de4e0%2Fkaggle3.png?generation=1675454640048750&alt=media =265x200)\n\nData Explorer does a lot of other things, like generate precision-recall curves, view images in confusion matrices, and perform coreset selection for faster training iteration. If you’re curious about my initial findings (and how to reproduce them), check out my [Data Explorer for Kaggle](https://akridata.ai/blog/data-explorer-for-kaggle-rsna/) blog.  Thanks!",
    "2131675": "Magic! Great tool and inspiration. I have spent a lot of time filtering data with pandas and ploting images in notebook to understand dataset.  This tool could be really very helpful. \"coreset\" option is my takeaway - learning from today 😁\n\nDo you have any standard projects (playground) where we can experiment without uploading the same data (rsna)?",
    "2132809": "Let me add a quick note on the [model analyze feature](https://docs.akridata.ai/docs/overview-analyze)... \n\nIf you have predictions from cross validation, upload them as a CSV file to get a confusion matrix view into the data. This is a convenient way to browse true positives vs. false positives, for example.\n\n**True positive predictions:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2Fb772edd28e96699792ec3b4253f8da26%2Frsna-true-pos.gif?generation=1675736993267472&alt=media)\n\n**False positive predictions:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7837%2F1be3bdeced0d38ef70781f79961e1adc%2Frsna-false-pos.gif?generation=1675737951242086&alt=media)",
    "2132682": "<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Manual inspection of data has probably the highest value-to-prestige ratio of any activity in machine learning.</p>&mdash; Greg Brockman (@gdb) <a href=\"https://twitter.com/gdb/status/1622683988736479232?ref_src=twsrc%5Etfw\">February 6, 2023</a></blockquote> <script async src=\"https://platform.twitter.com/widgets.js\" charset=\"utf-8\"></script>",
    "2131133": "After the initial visualization of the dataset, you can subsample it in various ways to train only on a portion of the data.\nFor example, the \"coreset\" option will preserves small clusters, but others, like the \"random\" sampling method are available too.",
    "2128702": "Thanks for promoting this, it is always good to find tools to enhance daily routines. Does it handle other kind of data too, or image data only?",
    "2131901": "kaggle is a platform for uploading data tool. kaggler earn their badges by creating data tool. If you can create some simple plugin that kaggler can use after they create their dataset, it will be good. kind of automatic EDA button at the data creation page. Now we have gptchat which can automatically generative data description ....\n\ni am late ... chatgpt EDA\nhttps://medium.com/@avra42/chatgpt-build-this-data-science-web-app-using-streamlit-python-25acca3cecd4"
  }
}