{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Still under construction \n\n# What is about ?\n\nLearn pytorch using CAFA5 as an example. Educational notebook 02. \n\nSee the previous notebook: https://www.kaggle.com/code/alexandervc/pytorch-01-basics for basics of pytorch - metrics, setting neural networks, training neural networks in the most simple setup. \n\nHere we  learn how to do some additional things:\n\n    + setup for cross-validation with 5 folds\n    + more complicated models - not Sequential, but with skip-connections and various activations, etc.. \n    + several examples of optimizers including external one -  e.g. \"Sophia\"\n    + several examples of losses - including external - focal, dice, Tversky, etc... \n    + several examples of learning rate schedulers e.g.  StepLR,  CosineAnnealingWarmRestarts \n\nWe also extend possibilities realted to the CAFA5 challenge:\n    \n    + one can choose among several embeddings available: t5, esm2 with different sizes  \n    + any number of targets - we load targets as sparse matrix with all 31466 train targets - one can choose how many targets one wants to predict, but be careful not to crash RAM - e.g. 3000-5000 targets might crash it. \n\nAnd more techinal thing: \n\n    + logging available memory - at key steps  we save to the log file - RAM available by  psutil. Because crashing by RAM happens often for that competition.  \n\n#### Educational task:\n\nChange neural network, lossses, optimizers, etc and try to improve the score. \n\n#### Links and Acknowledgements:\n\nWe owe a lot to (please upvote these  works):   \n\nModels used are borrowed from: \n\n\nLB 0.52875 Andrey (Sophia, skip-connections, etc) https://www.kaggle.com/code/andreylalaley/pytorch-cafa-5-prediction?scriptVersionId=138595845\n(quite sensitive to the choice of the random seed. Same model here gives 0.51643 See V36. See also experiments in https://www.kaggle.com/code/stasborodynkin/pytorch-cafa-5-prediction.  ) \n\nLB 0.49651 Ivan (Droupouts) https://www.kaggle.com/code/mrgobus/pytorch-01-basics-49cdbc?scriptVersionId=138589166&cellId=27\n\nLB 0.49545 Boris (PReLU) https://www.kaggle.com/code/sebastian157/pytorch-01-basics-3a6104?scriptVersionId=138569077&cellId=18\n( Same model here - 0.50277 - see V33)\n\nLB 0.4908 Lidia (LR scheduler - CosineAnnealingWarmRestarts, GELU)  https://www.kaggle.com/code/lidiashishina/pytorch-01-basics?scriptVersionId=138484544&cellId=18\n\n\nGreat collection of losses for pytorch and Keras: https://www.kaggle.com/code/bigironsphere/loss-function-library-keras-pytorch#Focal-Tversky-Loss\n\nFocal loss in pytorch can also be found here: https://pypi.org/project/focal-loss-pytorch/\n\nSchedulers for learning rate in pytorch https://www.kaggle.com/code/isbhargav/guide-to-pytorch-learning-rate-scheduling\n\nSee also: https://towardsdatascience.com/a-visual-guide-to-learning-rate-schedulers-in-pytorch-24bbb262c8\n\n\n\n#### Versions: \n    \n    58 Made public - 'quick_run' \n    36 LB 0.51643 - Andrey's model ","metadata":{}},{"cell_type":"markdown","source":"# Key params ","metadata":{}},{"cell_type":"code","source":"mode_loc = 'full' #  'quick_run'\nif mode_loc == 'quick_run':\n    # Crude downsampling - just for fast run/check - about 2-3 minutes for run \n    n_samples_to_consider =  10000#    downsampling - might be useful for debug:  small number - fast run \n    n_labels_to_consider = 10 # Up to 31466 but more than 3000-5000 may crash RAM \nelse:\n    n_samples_to_consider =   142246 #  1000 #   50_000#   142246 #    downsampling - might be useful for debug:  small number - fast run \n    n_labels_to_consider = 3000 # Up to 31466 but more than 3000-5000 may crash RAM \n\n# Choose emebeddings: \nfeatures_id = 't5'#   'esm2S2560' #  'esm2S1280' #  'esm2S320' #   \n\nmax_epoch = 10\nBATCH_SIZE = 128\nLEARNING_RATE = 0.005\nmode_submit = True \n\n# Loss_function for optimization:\nloss_criterion = {'loss':'BCE'}#  {'loss':'Focal'}#   {'loss':'Dice'}#    {'loss':'Tversky'}#  \n# Choose optimizer:\noptimizer_params = {'method': 'SophiaG'} #    {'method': 'Adam'} # {'method': 'RMSprop'} #    {'method': 'SGD'}\n# Choose learning rate scheduler\nscheduler_params = {'method': 'CosineAnnealingWarmRestarts' }  #  {'method':  'StepLR' } #  {'method':  None } #   \n\n\ncutoff_threshold_low = 0.01 # prediction < cutoff_threshold_low will be set to zero (i.e. no need to save to submission file)\n\nRANDOM_SEED = None # Fix or Not random seed \n\n\nstr_id = 'pMLP_'+features_id+'_Y'+str(n_labels_to_consider) + '_LR'+'%.0e'%(LEARNING_RATE)+'_BS'+str(BATCH_SIZE)\nstr_id += '_'+str(optimizer_params['method'] )\nstr_id += '_'+str(str(scheduler_params['method'])[:4] )\nstr_id += '_'+str(loss_criterion['loss'] )\nstr_id += '_S'+str(n_samples_to_consider)\nstr_id += '_CUT'+str(cutoff_threshold_low)\nprint(str_id, len(str_id))\nprint(str_id[:48])","metadata":{"execution":{"iopub.status.busy":"2023-08-03T13:16:23.139948Z","iopub.execute_input":"2023-08-03T13:16:23.140403Z","iopub.status.idle":"2023-08-03T13:16:23.151951Z","shell.execute_reply.started":"2023-08-03T13:16:23.140353Z","shell.execute_reply":"2023-08-03T13:16:23.150782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"logs_file_path = 'logs.txt'\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:45:48.3663Z","iopub.execute_input":"2023-08-03T09:45:48.366697Z","iopub.status.idle":"2023-08-03T09:45:48.385906Z","shell.execute_reply.started":"2023-08-03T09:45:48.366666Z","shell.execute_reply":"2023-08-03T09:45:48.384453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preparations \n\nTechnical installs/imports","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport time\nt0start = time.time()\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:45:48.387822Z","iopub.execute_input":"2023-08-03T09:45:48.388205Z","iopub.status.idle":"2023-08-03T09:45:48.423297Z","shell.execute_reply.started":"2023-08-03T09:45:48.388179Z","shell.execute_reply":"2023-08-03T09:45:48.422399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import psutil\nimport datetime\n\ndef get_available_ram():\n    virtual_memory = psutil.virtual_memory()\n    available_ram = virtual_memory.available\n    return available_ram\ndef log_available_ram( str_for_logging_optional = None):\n    try:\n        virtual_memory = psutil.virtual_memory()\n        available_ram_bytes = virtual_memory.available\n        # Convert bytes to other units if needed (e.g., megabytes, gigabytes)\n        available_ram_megabytes = available_ram_bytes / (1024 ** 2)\n        available_ram_gigabytes = available_ram_bytes / (1024 ** 3)\n\n    #     print(f\"Available RAM: {available_ram_bytes} bytes\")\n    #     print(f\"Available RAM: {available_ram_megabytes:.2f} MB\")\n        if str_for_logging_optional is not None:\n            print(str_for_logging_optional)\n        current_datetime = datetime.datetime.now()\n        str1 = f\"Available RAM: {available_ram_gigabytes:.2f} G  Current datetime: {current_datetime}\"\n        print( str1 ) \n        \n        with open(logs_file_path, 'a') as file:\n            if str_for_logging_optional is not None:\n                file.write(str_for_logging_optional + '\\n')\n            file.write(str1 + '\\n')\n        # print(\"Data appended successfully.\")\n    except Exception as e:\n        print(f\"Error while appending data: {e}\")        \n    \n#     return available_ram\nlog_available_ram('On start')\n\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:45:48.425262Z","iopub.execute_input":"2023-08-03T09:45:48.425709Z","iopub.status.idle":"2023-08-03T09:45:48.434777Z","shell.execute_reply.started":"2023-08-03T09:45:48.425671Z","shell.execute_reply":"2023-08-03T09:45:48.433498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture \n!pip install torchmetrics","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:45:48.437616Z","iopub.execute_input":"2023-08-03T09:45:48.438076Z","iopub.status.idle":"2023-08-03T09:45:57.434202Z","shell.execute_reply.started":"2023-08-03T09:45:48.438039Z","shell.execute_reply":"2023-08-03T09:45:57.43215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%capture\n!pip install torchsummary","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:45:57.436379Z","iopub.execute_input":"2023-08-03T09:45:57.43674Z","iopub.status.idle":"2023-08-03T09:46:06.588392Z","shell.execute_reply.started":"2023-08-03T09:45:57.436714Z","shell.execute_reply":"2023-08-03T09:46:06.586999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## https://www.kaggle.com/code/alexandervc/baseline-multilabel-to-multitarget-binary#Load-train-features---precalculated-embeddings-for-the-proteins\nimport os\nimport gc\nfrom sklearn.model_selection import train_test_split\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\ntqdm.pandas()\nimport torch\nimport warnings\nwarnings.filterwarnings('ignore')\nfrom sklearn.model_selection import train_test_split\nimport torch.nn.functional as F\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchmetrics import AUROC,F1Score\nfrom torchmetrics.classification import BinaryF1Score\nfrom torchsummary import summary as torchsummary\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-03T09:46:06.590694Z","iopub.execute_input":"2023-08-03T09:46:06.591017Z","iopub.status.idle":"2023-08-03T09:46:06.600104Z","shell.execute_reply.started":"2023-08-03T09:46:06.590988Z","shell.execute_reply":"2023-08-03T09:46:06.598978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nimport random \ndef seed_all(RANDOM_SEED):\n    if RANDOM_SEED is not None: \n        try:\n            SEED = RANDOM_SEED\n            random.seed(SEED)\n            np.random.seed(SEED)\n            torch.manual_seed(SEED)\n            torch.cuda.manual_seed_all(SEED)\n            torch.backends.cudnn.deterministic = True\n            torch.backends.cudnn.benchmark = False\n\n        except Exception as e:\n            print(f\"Exception: {e}\")\n            \nseed_all(RANDOM_SEED)            ","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:06.601813Z","iopub.execute_input":"2023-08-03T09:46:06.602183Z","iopub.status.idle":"2023-08-03T09:46:06.622176Z","shell.execute_reply.started":"2023-08-03T09:46:06.602155Z","shell.execute_reply":"2023-08-03T09:46:06.620672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Optimizer \"Sophia\" sometimes better than Adam","metadata":{}},{"cell_type":"code","source":"%%capture\n!git clone https://github.com/kyegomez/Sophia.git\n! python Sophia/setup.py install\n!rm Sophia/Sophia/__init__.py\nfrom Sophia.Sophia.Sophia import SophiaG","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:06.62358Z","iopub.execute_input":"2023-08-03T09:46:06.623871Z","iopub.status.idle":"2023-08-03T09:46:09.856185Z","shell.execute_reply.started":"2023-08-03T09:46:06.623848Z","shell.execute_reply":"2023-08-03T09:46:09.854573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Collection of custom losses \n\nBorrowed from: https://www.kaggle.com/code/bigironsphere/loss-function-library-keras-pytorch - one can find more losses there ","metadata":{}},{"cell_type":"markdown","source":"## Focal loss\n\nFocal Loss was introduced by Lin et al of Facebook AI Research in 2017 as a means of combatting extremely imbalanced datasets where positive cases were relatively rare. Their paper \"Focal Loss for Dense Object Detection\" is retrievable here: https://arxiv.org/abs/1708.02002. In practice, the researchers used an alpha-modified version of the function so I have included it in this implementation.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.8\nGAMMA = 2\n\nclass FocalLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(FocalLoss, self).__init__()\n\n    def forward(self, inputs, targets, alpha=ALPHA, gamma=GAMMA, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #first compute binary cross-entropy \n        BCE = F.binary_cross_entropy(inputs, targets, reduction='mean')\n        BCE_EXP = torch.exp(-BCE)\n        focal_loss = alpha * (1-BCE_EXP)**gamma * BCE\n                       \n        return focal_loss","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:09.858021Z","iopub.execute_input":"2023-08-03T09:46:09.859856Z","iopub.status.idle":"2023-08-03T09:46:09.868011Z","shell.execute_reply.started":"2023-08-03T09:46:09.859809Z","shell.execute_reply":"2023-08-03T09:46:09.866917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dice Loss","metadata":{}},{"cell_type":"code","source":"#PyTorch\nclass DiceLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(DiceLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        intersection = (inputs * targets).sum()                            \n        dice = (2.*intersection + smooth)/(inputs.sum() + targets.sum() + smooth)  \n        \n        return 1 - dice","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:09.86935Z","iopub.execute_input":"2023-08-03T09:46:09.86971Z","iopub.status.idle":"2023-08-03T09:46:09.89136Z","shell.execute_reply.started":"2023-08-03T09:46:09.869683Z","shell.execute_reply":"2023-08-03T09:46:09.889771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Tversky Loss\n\nThis loss was introduced in \"Tversky loss function for image segmentationusing 3D fully convolutional deep networks\", retrievable here: https://arxiv.org/abs/1706.05721. It was designed to optimise segmentation on imbalanced medical datasets by utilising constants that can adjust how harshly different types of error are penalised in the loss function. From the paper:\n\n... in the case of α=β=0.5 the Tversky index simplifies to be the same as the Dice coefficient, which is also equal to the F1 score. With α=β=1, Equation 2 produces Tanimoto coefficient, and setting α+β=1 produces the set of Fβ scores. Larger βs weigh recall higher than precision (by placing more emphasis on false negatives).\n\nTo summarise, this loss function is weighted by the constants 'alpha' and 'beta' that penalise false positives and false negatives respectively to a higher degree in the loss function as their value is increased. The beta constant in particular has applications in situations where models can obtain misleadingly positive performance via highly conservative prediction. You may want to experiment with different values to find the optimum. With alpha==beta==0.5, this loss becomes equivalent to Dice Loss.","metadata":{}},{"cell_type":"code","source":"#PyTorch\nALPHA = 0.5\nBETA = 0.5\n\nclass TverskyLoss(nn.Module):\n    def __init__(self, weight=None, size_average=True):\n        super(TverskyLoss, self).__init__()\n\n    def forward(self, inputs, targets, smooth=1, alpha=ALPHA, beta=BETA):\n        \n        #comment out if your model contains a sigmoid or equivalent activation layer\n        inputs = F.sigmoid(inputs)       \n        \n        #flatten label and prediction tensors\n        inputs = inputs.view(-1)\n        targets = targets.view(-1)\n        \n        #True Positives, False Positives & False Negatives\n        TP = (inputs * targets).sum()    \n        FP = ((1-targets) * inputs).sum()\n        FN = (targets * (1-inputs)).sum()\n       \n        Tversky = (TP + smooth) / (TP + alpha*FP + beta*FN + smooth)  \n        \n        return 1 - Tversky","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:09.89649Z","iopub.execute_input":"2023-08-03T09:46:09.896867Z","iopub.status.idle":"2023-08-03T09:46:09.90648Z","shell.execute_reply.started":"2023-08-03T09:46:09.896844Z","shell.execute_reply":"2023-08-03T09:46:09.905612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Choose device - GPU or CPU and assign model to it","metadata":{}},{"cell_type":"code","source":"device = torch.device(\"cuda:0\" if torch.cuda.is_available() else \"cpu\")\nprint(device)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:09.907686Z","iopub.execute_input":"2023-08-03T09:46:09.908007Z","iopub.status.idle":"2023-08-03T09:46:09.925991Z","shell.execute_reply.started":"2023-08-03T09:46:09.907977Z","shell.execute_reply":"2023-08-03T09:46:09.924218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load X,Y, etc \n\nLoad precalulcated features - embeddings and targets transformed in 0,1 multi-target task.\n\nWe use already computed emebdding for the proteins sequences - thanks to Andrey Shevtsov for sharing the embeddings by esm2-model: \n( https://www.kaggle.com/competitions/cafa-5-protein-function-prediction/discussion/406168 - please upvote his work ).\n\nWe load targets matrix \"Y\" which contains NOT all the CAFA5 targets but only top 1499. That is quite enough to get not so bad score. \nNote: original input file text describing targets have been trasnformed to np.array Y.\nSee the notebook https://www.kaggle.com/code/alexandervc/baseline-multilabel-to-multitarget-binary for details. \n","metadata":{}},{"cell_type":"code","source":"%%time \n\ndef get_paths_to_features(features_id):\n    \n    if features_id == 'esm2S2560':\n        fn_X = '/kaggle/input/4637427/train_embeds_esm2_t36_3B_UR50D.npy'\n        fn_prot_ids = '/kaggle/input/4637427/train_ids_esm2_t36_3B_UR50D.npy'\n        fn_X_submit = '/kaggle/input/4637427/test_embeds_esm2_t36_3B_UR50D.npy'\n        fn_submit_protein_ids = '/kaggle/input/4637427/test_ids_esm2_t36_3B_UR50D.npy'\n        \n    elif features_id == 'esm2S1280':\n        fn_X = '/kaggle/input/23468234/train_embeds_esm2_t33_650M_UR50D.npy'\n        fn_prot_ids = '/kaggle/input/23468234/train_ids_esm2_t33_650M_UR50D.npy'\n        fn_X_submit = '/kaggle/input/23468234/test_embeds_esm2_t33_650M_UR50D.npy'\n        fn_submit_protein_ids = '/kaggle/input/23468234/test_ids_esm2_t33_650M_UR50D.npy'\n        \n    elif features_id == 't5':\n        fn_X = '/kaggle/input/cafa5-features-etc/T5_train_embeds_float32.npy'\n        fn_prot_ids = '/kaggle/input/cafa5-features-etc/train_ids.npy'\n        fn_X_submit = '/kaggle/input/cafa5-features-etc/T5_test_embeds_float32.npy'\n        fn_submit_protein_ids = '/kaggle/input/cafa5-features-etc/test_ids.npy'\n        \n    elif features_id == 'esm2S320':\n        fn_X = '/kaggle/input/315701375/train_embeds_esm2_t6_8M_UR50D.npy'\n        fn_prot_ids = '/kaggle/input/315701375/train_ids_esm2_t6_8M_UR50D.npy'\n        fn_X_submit = '/kaggle/input/315701375/test_embeds_esm2_t6_8M_UR50D.npy'\n        fn_submit_protein_ids = '/kaggle/input/315701375/test_ids_esm2_t6_8M_UR50D.npy'\n\n    return fn_X, fn_prot_ids, fn_X_submit, fn_submit_protein_ids \n    \n################ load  features #########################\nfn_X, fn_prot_ids, fn_X_submit, fn_submit_protein_ids  = get_paths_to_features(features_id)\n\nfn = fn_X #  '/kaggle/input/4637427/train_embeds_esm2_t36_3B_UR50D.npy'\nX = np.load(fn).astype(np.float32)[:n_samples_to_consider, :]\nprint(X.shape)\nprint(X[:2,:3])\nprot_ids  = np.load(fn_prot_ids)\nprint(prot_ids.shape)\nprint(prot_ids[:15])\n\n################ load  features for submit #########################\nif mode_submit:\n    # fn = '/kaggle/input/4637427/train_embeds_esm2_t36_3B_UR50D.npy'\n    # fn = '/kaggle/input/4637427/test_embeds_esm2_t36_3B_UR50D.npy'\n    fn = fn_X_submit\n    print(fn)\n    # X_submit = np.load(fn).astype(np.float32)\n    X_submit = torch.tensor(np.load(fn), dtype=torch.float32).to(device)\n    print(X_submit.shape)\n    print(X_submit[:2,:3])\n\n    fn = fn_submit_protein_ids\n    submit_protein_ids = np.load(fn)\n    print(submit_protein_ids.shape, submit_protein_ids[:10])\n    \n\n############################ load targets and their ids  ######################################\n\n# Load targets Y\nfrom scipy import sparse\nfn = '/kaggle/input/cafa5-features-etc/Y_31466_sparse_float32.npz'\nY = sparse.load_npz(fn )\nprint('Y', Y.shape, 'loaded')\nY = Y[:n_samples_to_consider,:n_labels_to_consider].toarray()\nprint('Y', Y.shape, 'truncated')\nn_labels_to_consider = Y.shape[1] # in case n_labels_to_consider is greater that Y.shape we decrease it\n\nfn = '/kaggle/input/cafa5-features-etc/Y_31466_labels.npy'\nY_labels = np.load(fn, allow_pickle=True )[:n_labels_to_consider]\nprint(Y_labels.shape)\nprint(Y_labels[:20])\n\n# %%time\nif 1:\n    v = Y.sum(axis = 0)\n    plt.figure(figsize = (20,6))\n    plt.plot(v[:500], '*-')\n    plt.grid()\n    plt.title(' Number of 1 in targets',fontsize = 20 )\n    plt.xlabel('target index', fontsize = 20 )\n    plt.show()\n\n    \n\n\n# fn_train_terms = '/kaggle/input/cafa-5-protein-function-prediction/Train/train_terms.tsv'\n# fn_train_taxonomy ='/kaggle/input/cafa-5-protein-function-prediction/Train/train_taxonomy.tsv'\n\nimport gc\ngc.collect()\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:09.927807Z","iopub.execute_input":"2023-08-03T09:46:09.928182Z","iopub.status.idle":"2023-08-03T09:46:11.151852Z","shell.execute_reply.started":"2023-08-03T09:46:09.928158Z","shell.execute_reply":"2023-08-03T09:46:11.150821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"log_available_ram('After data load')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.153144Z","iopub.execute_input":"2023-08-03T09:46:11.153777Z","iopub.status.idle":"2023-08-03T09:46:11.161896Z","shell.execute_reply.started":"2023-08-03T09:46:11.153747Z","shell.execute_reply":"2023-08-03T09:46:11.159969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load partition to folds for cross-validation\n\nfolds_gkf - 5-folds random split","metadata":{}},{"cell_type":"code","source":"%%time \nfn = '/kaggle/input/cafa5-features-etc/random_folds/folds_gkf.npy'\nfolds = np.load(fn)[:n_samples_to_consider]\nprint(folds.shape, len(set(folds)))\nfor k in set(folds):\n    m = folds == k\n    print(k, m.sum() )\nfolds\n\n\n# from sklearn.model_selection import train_test_split\n# IX_train,IX_val = train_test_split( np.arange(len(X)), train_size=0.7, random_state=42)\n# print(IX_train.shape,IX_val.shape) \n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.163762Z","iopub.execute_input":"2023-08-03T09:46:11.164078Z","iopub.status.idle":"2023-08-03T09:46:11.185158Z","shell.execute_reply.started":"2023-08-03T09:46:11.164053Z","shell.execute_reply":"2023-08-03T09:46:11.184125Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define neural network \n\nExample of simple neural network on pytorch. \n\n\nJust an example randomly taken from previous Kaggle comeptitions, nothing have been optmized for CAFA5 - you are welcome to do it. \n\n\n","metadata":{}},{"cell_type":"code","source":"import torch.nn as nn","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.186121Z","iopub.execute_input":"2023-08-03T09:46:11.1864Z","iopub.status.idle":"2023-08-03T09:46:11.198255Z","shell.execute_reply.started":"2023-08-03T09:46:11.186375Z","shell.execute_reply":"2023-08-03T09:46:11.197176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## https://www.kaggle.com/code/andreylalaley/pytorch-cafa-5-prediction?scriptVersionId=138595845\n## 0.52875 Andrey \nclass Model5(nn.Module):\n    def __init__(self,input_features,output_features):\n        super().__init__()\n        \n        self.activation = nn.PReLU()\n        \n        self.bn1 = nn.BatchNorm1d(input_features)\n        self.fc1 = nn.Linear(input_features, 800)\n        self.ln1 = nn.LayerNorm(800, elementwise_affine=True)\n        \n        self.bn2 = nn.BatchNorm1d(800)\n        self.fc2 = nn.Linear(800, 600)\n        self.ln2 = nn.LayerNorm(600, elementwise_affine=True)\n        \n        self.bn3 = nn.BatchNorm1d(600)\n        self.fc3 = nn.Linear(600, 400)\n        self.ln3 = nn.LayerNorm(400, elementwise_affine=True)\n        \n        self.bn4 = nn.BatchNorm1d(1200)\n        self.fc4 = nn.Linear(1200, output_features)\n        self.ln4 = nn.LayerNorm(output_features, elementwise_affine=True)\n        \n        self.sigm = nn.Sigmoid()\n    def forward(self,inputs):\n#         print(inputs.shape)\n\n        fc1_out = self.bn1(inputs)\n        fc1_out = self.ln1(self.fc1(inputs))\n        fc1_out = self.activation(fc1_out)\n        \n        x = self.bn2(fc1_out)\n        \n        x = self.ln2(self.fc2(x))\n        x = self.activation(x)\n        \n        x = self.bn3(x)\n        \n        x = self.ln3(self.fc3(x))\n        x = self.activation(x)\n        \n        x = torch.cat([x, fc1_out], axis = -1)\n        \n        x = self.bn4(x)\n        \n        x = self.ln4(self.fc4(x))\n        out = self.sigm(x)\n        return out\n    \n    \n## https://www.kaggle.com/code/mrgobus/pytorch-01-basics-49cdbc?scriptVersionId=138589166&cellId=27\n## 0.49651 Ivan\nclass Model4(nn.Module):\n\n    def __init__(self, in_features, out_features, neurons_per_layer = 1000):\n\n        super().__init__()\n\n        self.in_features = in_features\n        self.out_features = out_features\n\n        self.model = nn.Sequential(\n            nn.BatchNorm1d(in_features),\n            nn.Dropout(0.2),\n            nn.Linear(in_features, neurons_per_layer),\n            nn.LayerNorm(neurons_per_layer),\n            nn.PReLU(init = 0.5),\n\n            nn.BatchNorm1d(neurons_per_layer),\n            nn.Dropout(0.2),\n            nn.Linear(neurons_per_layer, neurons_per_layer),\n            nn.LayerNorm(neurons_per_layer),\n            nn.PReLU(init = 0.5),\n\n            nn.BatchNorm1d(neurons_per_layer),\n            nn.Linear(neurons_per_layer, out_features),\n            nn.LayerNorm(out_features),\n            nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        return self.model(x)\n    \n    \n## https://www.kaggle.com/code/sebastian157/pytorch-01-basics-3a6104?scriptVersionId=138569077&cellId=18\n## LB 0.49545  \nclass Model3(nn.Module):\n    def __init__(self, in_features, out_features):\n        super().__init__()\n        self.in_features = in_features\n        self.out_features = out_features\n\n        self.model = nn.Sequential(\n            nn.BatchNorm1d(in_features),\n         \n            nn.Linear(in_features, 800),\n            nn.LayerNorm(800, elementwise_affine=True),\n            nn.PReLU(),\n            \n            nn.BatchNorm1d(800),  \n            \n            nn.Linear(800, 600),\n            nn.LayerNorm(600, elementwise_affine=True),\n            nn.PReLU(),\n            \n            nn.BatchNorm1d(600),\n          \n            nn.Linear(600, 400),\n            nn.LayerNorm(400, elementwise_affine=True),\n            nn.PReLU(),\n            \n            nn.BatchNorm1d(400),\n       \n            nn.Linear(400, out_features),\n            nn.LayerNorm(out_features, elementwise_affine=True),\n            nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        return self.model(x)\n\n\n## https://www.kaggle.com/code/lidiashishina/pytorch-01-basics?scriptVersionId=138484544&cellId=18\n# LB 0.4908 Boris\nclass Model2(nn.Module):\n    def __init__(self, in_features, out_features):\n        super().__init__()\n        self.in_features = in_features\n        self.out_features = out_features\n\n        self.model = nn.Sequential(\n            nn.BatchNorm1d(in_features),\n         \n            nn.Linear(in_features, 800),\n            nn.LayerNorm(800, elementwise_affine=True),\n            nn.GELU(),\n      \n            nn.BatchNorm1d(800),  \n            \n            nn.Linear(800, 600),\n            nn.LayerNorm(600, elementwise_affine=True),\n            nn.GELU(),\n\n            nn.BatchNorm1d(600),\n          \n            nn.Linear(600, 400),\n            nn.LayerNorm(400, elementwise_affine=True),\n            nn.GELU(),\n\n            nn.BatchNorm1d(400),\n       \n            nn.Linear(400, out_features),\n            nn.LayerNorm(out_features, elementwise_affine=True),\n            nn.Sigmoid()\n        )\n\n    def forward(self, x):\n        return self.model(x)\n    \n# https://www.kaggle.com/code/alexandervc/moa-nn-04/notebook\n\nclass Model1(nn.Module): # The parent class for the models is nn.Module \n    \n    def __init__(self, in_features, out_features): # constructor \n        \n        super().__init__() # the constructor of the upper class is first called\n\n        self.in_features = in_features\n        self.out_features = out_features\n\n        self.model = nn.Sequential( #  Sequential addition of layers -  multi-layer perceptron \n            nn.BatchNorm1d(in_features),\n            nn.Linear(in_features, 1000),\n            nn.GELU(), # nn.ReLU(),\n\n            nn.BatchNorm1d(1000),            # \n            nn.Dropout(0.35),\n            nn.Linear(1000, 2000),\n            nn.GELU(),# nn.ReLU(),\n\n#             nn.BatchNorm1d(600),            # nn.Dropout(0.1),\n#             nn.Linear(600, 400),\n#             nn.ReLU(),\n\n#             nn.BatchNorm1d(400),\n            nn.Linear(2000, out_features),\n            nn.Sigmoid()\n        )\n\n    def forward(self, x): # \n        return self.model(x)","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.199287Z","iopub.execute_input":"2023-08-03T09:46:11.199613Z","iopub.status.idle":"2023-08-03T09:46:11.223416Z","shell.execute_reply.started":"2023-08-03T09:46:11.199571Z","shell.execute_reply":"2023-08-03T09:46:11.222572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Get model function","metadata":{}},{"cell_type":"code","source":"def get_model():\n    model = Model5(X.shape[1],Y.shape[1])\n    \n    model.to(device)\n    return model\n\nmodel = get_model()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.22453Z","iopub.execute_input":"2023-08-03T09:46:11.224823Z","iopub.status.idle":"2023-08-03T09:46:11.260336Z","shell.execute_reply.started":"2023-08-03T09:46:11.224801Z","shell.execute_reply":"2023-08-03T09:46:11.259441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# That does not work for us - why ? \ntry:\n    torchsummary(model, XX.size(), batch_size=-1, device= 'cuda')\nexcept    Exception as inst:\n    print(type(inst))    # the exception type\n    print(inst.args)     # arguments stored in .args\n    print(inst)          # __str__ allows args to be printed directly,\n                         # but may be overridden in exception subclasses\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.261397Z","iopub.execute_input":"2023-08-03T09:46:11.26232Z","iopub.status.idle":"2023-08-03T09:46:11.267811Z","shell.execute_reply.started":"2023-08-03T09:46:11.262292Z","shell.execute_reply":"2023-08-03T09:46:11.266541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Auxilliary functions compute/save scores, etc ","metadata":{}},{"cell_type":"code","source":"%%time\nimport time \nfrom torchmetrics.classification import BinaryF1Score\nfrom torchmetrics import AUROC\nfrom torch.utils.data import DataLoader, TensorDataset\n\ndef update_modeling_stat(df_stat,str_id, model, X_train, Y_train, X_val, Y_val, IX_train, IX_val, \n                         epoch , ix_fold, t0_epoch_train, dict_optional_info = {}, record_index = None, verbose = 0  ):\n    '''\n    Compute/store/save scores/metrics/statistics on modelling.\n    '''\n    if verbose >= 100:\n        print('Scoring starts')\n        \n    IX_df_stat = record_index\n    if IX_df_stat is None:\n        IX_df_stat = len(df_stat)+1\n    \n    \n    t0 = time.time()\n    df_stat.loc[IX_df_stat,'Id'] = str_id\n    df_stat.loc[IX_df_stat,'epoch'] = epoch\n    df_stat.loc[IX_df_stat,'fold'] = ix_fold\n    \n    model.eval()\n    with torch.no_grad():\n        preds_val = model(X_val)\n        if X_train is not None: \n            preds_train = model(X_train)\n        \n        df_stat.loc[IX_df_stat,'Loss Val'] = criterion(preds_val, Y_val).item()\n        if X_train is not None: \n            df_stat.loc[IX_df_stat,'Loss Train'] = criterion(preds_train, Y_train).item()\n        \n        #from torchmetrics.classification import BinaryF1Score\n        f1 = BinaryF1Score(threshold=0.25, multidim_average= 'samplewise'  )\n        s = f1(preds_val, Y_val) # vector of values for each sample\n        df_stat.loc[IX_df_stat, 'F1|0.25 Val'] = s.mean().item()#  average over samples\n        if X_train is not None: \n            s = f1(preds_train, Y_train)  \n            df_stat.loc[IX_df_stat, 'F1|0.25 Train'] =  s.mean().item()        \n        \n        \n        #from torchmetrics import AUROC\n        auroc = AUROC(task = 'binary')\n        s = auroc(preds_val, Y_val) # auroc between flattened arguments \n        df_stat.loc[IX_df_stat, 'AUC Val'] = s.item()\n        if X_train is not None: \n            s = auroc(preds_train, Y_train) # auroc between flattened arguments \n            df_stat.loc[IX_df_stat, 'AUC Train'] =  s.item()        \n\n        \n        for threshold_for_f1 in [0.2,0.3]:\n            #from torchmetrics.classification import BinaryF1Score\n            f1 = BinaryF1Score(threshold=threshold_for_f1, multidim_average= 'samplewise'  )\n            s = f1(preds_val, Y_val) # vector of values for each sample\n            df_stat.loc[IX_df_stat, 'F1|'+str(threshold_for_f1)+' Val'] = s.mean().item()#  average over samples\n            if X_train is not None: \n                s = f1(preds_train, Y_train) \n                df_stat.loc[IX_df_stat, 'F1|'+str(threshold_for_f1)+' Train'] =  s.mean().item()\n            \n            \n    df_stat.loc[IX_df_stat, 'n_targets'] = Y_val.shape[1]\n    df_stat.loc[IX_df_stat, 'n_samples Val'] = Y_val.shape[0]\n    df_stat.loc[IX_df_stat, 'Time scoring'] = np.round(time.time() - t0, 1 )    \n    df_stat.loc[IX_df_stat, 'Time train epoch'] = np.round(t0_epoch_train , 1 )\n    df_stat.loc[IX_df_stat, 'Time scoring'] = np.round(time.time() - t0, 1 )    \n    for k in dict_optional_info:\n        val = dict_optional_info[k]\n        df_stat.loc[IX_df_stat, k] = val   \n        \n    df_stat.to_csv('df_foldwise_stat.csv')        \n    if verbose > 0:\n        display(df_stat.tail(1))\n        print('Scoring finished. Seconds passed:  %.1f'%(time.time() - t0)  )\n\n           \n    return df_stat\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.268834Z","iopub.execute_input":"2023-08-03T09:46:11.269311Z","iopub.status.idle":"2023-08-03T09:46:11.299882Z","shell.execute_reply.started":"2023-08-03T09:46:11.269287Z","shell.execute_reply":"2023-08-03T09:46:11.298628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train the model","metadata":{}},{"cell_type":"code","source":"%%time\n# import datetime\n# current_datetime = datetime.datetime.now()\n# print(\"Current datetime:\", current_datetime)\nlog_available_ram('Before modeling')\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.30142Z","iopub.execute_input":"2023-08-03T09:46:11.30266Z","iopub.status.idle":"2023-08-03T09:46:11.320097Z","shell.execute_reply.started":"2023-08-03T09:46:11.302624Z","shell.execute_reply":"2023-08-03T09:46:11.319128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n######################### Params ##################################################3\n\nif 0: # Params setting moved to the top of the notebook - section \"Key params\"\n    max_epoch = 10\n    BATCH_SIZE = 128\n    LEARNING_RATE = 0.001\n    mode_submit = True \nverbose = 1000\n\n######################### Output ##################################################\ndf_stat = pd.DataFrame()\n\n########################## Preparations ###########################################\n\n#         ######### import etc  ######\nfrom torchmetrics.classification import BinaryF1Score\nfrom torchmetrics import AUROC\nimport gc \nimport time \n\nprint(str_id)\nprint();print();\n\nif loss_criterion['loss'] == 'BCE':\n    criterion = nn.BCELoss()\nelif loss_criterion['loss'] == 'Focal':\n    criterion =  FocalLoss()\nelif loss_criterion['loss'] == 'Dice':\n    criterion =  DiceLoss()\nelif loss_criterion['loss'] == 'Tversky':\n    criterion =   TverskyLoss()\nprint( loss_criterion, str( criterion ) )\n\n\n#         #########  Other stuff to initialize ######\n\ncnt_blend_submit = 0 ; \nif (mode_submit is not None) and ( mode_submit != False ): # None\n    Y_submit = np.zeros( (141865, Y.shape[1] )  , dtype = np.float16 ) \n    \nt0 = time.time()\nprint(); print('Start training NN',datetime.datetime.now()) ; print()\nlist_folds_ix =  np.sort(list ( set(folds))  )\n\nlog_available_ram('Right before modeling')\n\n########################## Main modelling  ###########################################\nfor ix_fold  in  list_folds_ix:\n    \n    model = get_model() # Reinitialize model each time, otherwise weights will be updated\n    \n    if optimizer_params['method'] == 'Adam':\n        optimizer = torch.optim.Adam(model.parameters(), lr=LEARNING_RATE) \n    elif optimizer_params['method'] == 'SophiaG':\n        optimizer = SophiaG(model.parameters(),lr=LEARNING_RATE, betas=(0.965, 0.99), rho = 0.01, weight_decay=1e-1)\n    elif optimizer_params['method'] == 'RMSprop':\n        optimizer = torch.optim.RMSprop(model.parameters(), lr=LEARNING_RATE) \n    elif optimizer_params['method'] == 'SGD':\n        optimizer = torch.optim.SGD(model.parameters(), lr=LEARNING_RATE)\n\n    if scheduler_params['method'] is None:\n        lr_sched = None \n    elif scheduler_params['method'] == 'CosineAnnealingWarmRestarts':\n        lr_sched = torch.optim.lr_scheduler.CosineAnnealingWarmRestarts(optimizer, T_0=10, T_mult=2, eta_min=0.001, last_epoch=-1) # https://www.kaggle.com/code/lidiashishina/pytorch-01-basics?scriptVersionId=138484544&cellId=33\n    elif scheduler_params['method'] == 'StepLR':\n        lr_sched = torch.optim.lr_scheduler.StepLR(optimizer, 0.5, 5)\n    \n    \n    ##################### Prepare train data ###################################################\n    mask_fold = folds == ix_fold\n    IX_train = np.where(mask_fold ==  0)[0]; IX_val = np.where(mask_fold > 0 )[0]; \n    X_train = torch.tensor(X[IX_train,:], dtype=torch.float32).to(device)\n    Y_train = torch.tensor(Y[IX_train,:], dtype=torch.float32).to(device)\n    X_val = torch.tensor(X[IX_val,:], dtype=torch.float32).to(device)\n    Y_val = torch.tensor(Y[IX_val,:], dtype=torch.float32).to(device)\n    train_dataset = TensorDataset(X_train, Y_train)\n    train_dataloader = DataLoader(train_dataset, batch_size=BATCH_SIZE, shuffle=True) \n    if verbose >= 100:\n        print(model)\n        print('fold',ix_fold, 'count:', mask_fold.sum(), 'n_targets:', Y_train.shape[1], 'Time: %.1f'%( time.time() - t0  )  )\n        print('X_train.shape, Y_train.shape', X_train.shape, Y_train.shape); print();\n        \n    log_available_ram(f'Right After Model, X_train/val initialization. Training loop fold: {ix_fold}')        \n    \n    ##################### Model Training ###################################################\n    for epoch in range(max_epoch):\n        t0_epoch = time.time()\n        model.train() # switch model into train mode. This helps inform layers such as Dropout and BatchNorm, which are designed to behave differently during training and evaluation. For instance, in training mode, BatchNorm updates a moving average on each new batch; whereas, for evaluation mode, these updates are frozen.\n        for i_batch, (x_batch, y_batch) in enumerate(train_dataloader): # Loop ove batches\n#             x_batch, y_batch = x_batch.to(device), y_batch.to(device) # do we need it ? may be already on device \n            optimizer.zero_grad() # technical - set gradients to zero, otherwise they will be accumulated         \n    \n            preds = model(x_batch)# Compute predictions only for batch samples \n        \n            loss = criterion(preds, y_batch) # Compute loss function for batch predictions\n            \n            loss.backward() # Compute gradients\n            optimizer.step() # Update NN weights using gradients\n            \n        if lr_sched is not None:\n            lr_sched.step() # Step LR scheduler\n                \n        if (verbose >= 10 ) and (i_batch % 100 == 0)  :\n            print(f'fold: {ix_fold}, Epoch: {epoch}, batch: {i_batch},  train loss on batch: {loss.item():12.5f} , time: {time.time() - t0:.1f} ' )            \n\n    \n        ##################### Compute/store scores,metrics,statistics, etc ###################################################\n        t0_epoch_train = time.time() - t0_epoch\n        update_modeling_stat(df_stat,str_id, model, X_train, Y_train, X_val, Y_val, IX_train, IX_val, \n                             epoch , ix_fold, t0_epoch_train, dict_optional_info = {}, record_index = None, verbose = 1000000   )\n    \n    ##################### Prepare predict for submission ###################################################\n    if  (mode_submit is not None) and ( mode_submit != False ):    \n        t0_submit = time.time()\n        model.eval()\n        with torch.no_grad():\n            Y_submit = (Y_submit * cnt_blend_submit  + model(X_submit).numpy() )/ (cnt_blend_submit + 1);  # Average predictions from different folds/models\n            cnt_blend_submit += 1 \n        if verbose >= 10: print(f'predict on submit time: {time.time() - t0_submit :.1f} ' )  \n            \n\n    del X_train, Y_train, X_val, Y_val,  train_dataset, train_dataloader\n    gc.collect()\n    torch.cuda.empty_cache()\n    if verbose > 0:\n        log_available_ram(f'Training loop fold: {ix_fold}')        \n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:46:11.321284Z","iopub.execute_input":"2023-08-03T09:46:11.321578Z","iopub.status.idle":"2023-08-03T09:48:08.752848Z","shell.execute_reply.started":"2023-08-03T09:46:11.321538Z","shell.execute_reply":"2023-08-03T09:48:08.751628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_stat.shape)\ndisplay( df_stat )","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:08.754156Z","iopub.execute_input":"2023-08-03T09:48:08.754448Z","iopub.status.idle":"2023-08-03T09:48:08.804752Z","shell.execute_reply.started":"2023-08-03T09:48:08.754424Z","shell.execute_reply":"2023-08-03T09:48:08.803664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Look on scoring results ","metadata":{}},{"cell_type":"code","source":"m = df_stat['epoch'] == df_stat['epoch'].max()\nprint('Final epoch stat')\ndisplay(df_stat[m])\ndf_stat[m].to_csv('final_epoch_df_stat.csv')\n\nprint('Final epoch foldwise mean:')\nd3 = df_stat[m].mean().to_frame().T\ndisplay( d3 )\nd3.to_csv('final_scores_only.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:08.806051Z","iopub.execute_input":"2023-08-03T09:48:08.806444Z","iopub.status.idle":"2023-08-03T09:48:08.850807Z","shell.execute_reply.started":"2023-08-03T09:48:08.80641Z","shell.execute_reply":"2023-08-03T09:48:08.849169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scores plots","metadata":{}},{"cell_type":"markdown","source":"## Epochwise mean over folds","metadata":{}},{"cell_type":"code","source":"d2 = df_stat.groupby('epoch').mean() \n\nlist_metrics_keywords = ['Loss', 'AUC']\n\nplt.figure(figsize = (20,8) )\nplt.suptitle('Mean over folds',fontsize = 20 )\nfor i,kw in enumerate(list_metrics_keywords):\n    l = [col for col in df_stat if kw in col]\n    plt.subplot(1,len(list_metrics_keywords ),i+1)\n    plt.plot( d2[l], label = l )\n    plt.grid()\n    plt.legend(fontsize = 12)\n    plt.title(kw, fontsize = 20)\n    plt.xlabel('epoch',fontsize = 15)\n    \nplt.show() \n\nlist_metrics_keywords = ['F1|0.25','F1|0.2', 'F1|0.3' ]\n\nplt.figure(figsize = (20,8) )\nplt.suptitle('Mean over folds',fontsize = 20 )\nfor i,kw in enumerate(list_metrics_keywords):\n    l = [col for col in df_stat if kw in col]\n    plt.subplot(1,len(list_metrics_keywords ),i+1)\n    plt.plot( d2[l], label = l )\n    plt.grid()\n    plt.legend(fontsize = 12)\n    plt.title(kw, fontsize = 20)\n    plt.xlabel('epoch',fontsize = 15)\n    \nplt.show() ","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:08.851959Z","iopub.execute_input":"2023-08-03T09:48:08.852295Z","iopub.status.idle":"2023-08-03T09:48:09.786215Z","shell.execute_reply.started":"2023-08-03T09:48:08.852268Z","shell.execute_reply":"2023-08-03T09:48:09.78463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Epochwise separate for each fold","metadata":{}},{"cell_type":"code","source":"\nfor ix_fold in np.sort( df_stat['fold'].unique()):\n    list_metrics_keywords = ['Loss', 'AUC']\n\n    print(ix_fold)\n    mask = df_stat['fold'] == ix_fold\n    print(mask.sum())\n    plt.figure(figsize = (20,8) )\n    plt.suptitle('Fold: '+str(ix_fold) , fontsize = 20 )\n    for i,kw in enumerate(list_metrics_keywords):\n        l = [col for col in df_stat if kw in col]\n        plt.subplot(1,len(list_metrics_keywords ),i+1)\n        plt.plot( df_stat[l][mask].values, label = l )\n        plt.grid()\n        plt.legend(fontsize = 12)\n        plt.title(kw, fontsize = 20)\n        plt.xlabel('epoch',fontsize = 15)\n\n    plt.show() \n\n    list_metrics_keywords = ['F1|0.25','F1|0.2', 'F1|0.3' ]\n\n    plt.figure(figsize = (20,8) )\n    plt.suptitle('Fold: '+str(ix_fold) , fontsize = 20 )\n    for i,kw in enumerate(list_metrics_keywords):\n        l = [col for col in df_stat if kw in col]\n        plt.subplot(1,len(list_metrics_keywords ),i+1)\n        plt.plot( df_stat[l][mask].values, label = l )\n        plt.grid()\n        plt.legend(fontsize = 12)\n        plt.title(kw, fontsize = 20)\n        plt.xlabel('epoch',fontsize = 15)\n\n    plt.show() \n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:09.792149Z","iopub.execute_input":"2023-08-03T09:48:09.792505Z","iopub.status.idle":"2023-08-03T09:48:14.305193Z","shell.execute_reply.started":"2023-08-03T09:48:09.792482Z","shell.execute_reply":"2023-08-03T09:48:14.30404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare submission","metadata":{}},{"cell_type":"code","source":"%%time\nimport gc\nif 0:\n    del X_train,Y_train, X_val, Y_val, train_dataset, train_dataloader, X, Y\ngc.collect()\ntorch.cuda.empty_cache()\n\nlog_available_ram()","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:14.306732Z","iopub.execute_input":"2023-08-03T09:48:14.307129Z","iopub.status.idle":"2023-08-03T09:48:14.760241Z","shell.execute_reply.started":"2023-08-03T09:48:14.307097Z","shell.execute_reply":"2023-08-03T09:48:14.758701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prepare submission tsv file","metadata":{}},{"cell_type":"code","source":"%%time\n\nmode_submit_prepare = 'slow_less_RAM_consuming'\n# 'slow_less_RAM_consuming' - works slower but consumes less RAM\n\nimport time \nt0 = time.time()\n\nprint( mode_submit_prepare , mode_submit)\nif mode_submit:\n    print(Y_submit.shape)\n    if mode_submit_prepare == 'slow_less_RAM_consuming':\n\n\n        file_path = \"submission.tsv\"\n        cc = 0\n        cc2 = 0\n        with open(file_path, 'w') as file:\n            for i in range(Y_submit.shape[0]):\n                for j in range(Y_submit.shape[1]):\n                    val = Y_submit[i,j]\n                    if val >= cutoff_threshold_low:\n                        str_go_term = str(Y_labels[j])\n                        str_protein_id = str( submit_protein_ids[i] )\n                        str_save = str_protein_id+'\\t'+str_go_term + '\\t' + '%.3f'%val + '\\n'\n                        file.write(str_save)   \n                        cc2 +=1\n                        if cc2 <= 10:\n                            if cc2 == 1: print('First 10 examples of the saved data:')\n                            print(str_save)\n                    cc += 1\n                    if cc % 30_000_000  == 0: \n                        sz = Y_submit.shape[0]*Y_submit.shape[1]\n                        print(cc, 'out of',sz, 'percent %.2f'%(cc/sz*100), 'saved:'  ,cc2, 'time %.1f'%(time.time() - t0 ))\n\n        print(cc2,'results saved to submission file', 'time  %.1f'%(time.time() - t0 )  )                \n\n    else:\n\n        # That is widely used way to preparase submission , but it might crash by RAM \n\n        df_submission = pd.DataFrame(columns = ['Protein Id', 'GO Term Id','Prediction'])\n\n        n_targets_predicted = Y_submit.shape[1]\n        n_samples_predicted = Y_submit.shape[0]\n        print('n_samples_predicted, n_targets_predicted',  n_samples_predicted, n_targets_predicted )\n\n\n        protein_list = []\n        for k in list(submit_protein_ids):\n            protein_list += [k] * n_targets_predicted\n        df_submission['Protein Id'] = protein_list\n\n        df_submission['GO Term Id'] = list(Y_labels) * n_samples_predicted\n        df_submission['Prediction'] = Y_submit.ravel()\n\n        df_submission = df_submission.round(3)\n        df_submission = df_submission[ df_submission['Prediction'] >= cutoff_threshold_low  ]\n\n        memory_usage_per_column = df_submission.memory_usage(deep=True)\n        total_memory_usage = memory_usage_per_column.sum()\n        print(\"\\nTotal memory usage:\", total_memory_usage/1e6, \"Megabytes\")\n\n        print(df_submission.shape)\n        display(df_submission)\n\n        import gc\n        if 0:\n            del preds \n\n        gc.collect()\n\n        df_submission.to_csv(\"submission.tsv\",header=False, index=False,sep='\\t')\n\n\n    log_available_ram('After saving submission')    ","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:14.761989Z","iopub.execute_input":"2023-08-03T09:48:14.762312Z","iopub.status.idle":"2023-08-03T09:48:19.607703Z","shell.execute_reply.started":"2023-08-03T09:48:14.762284Z","shell.execute_reply":"2023-08-03T09:48:19.606803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Plot histograms on submission predictions","metadata":{}},{"cell_type":"code","source":"%%time\ntry:\n    print(df_submission.shape)\n    plt.figure(figsize = (15,4))\n    plt.hist(df_submission['Prediction'].values, bins = 1000 )\n    plt.show()\n    print(df_submission.shape)\n    display(df_submission.describe())\n\n    for t in [0.1,0.2, 0.25,0.28, 0.3,0.4,0.5,0.6,0.7,0.8,0.9,1]:\n        m = df_submission['Prediction'] > t\n        print(t, m.sum(), m.sum()/ (n_samples_predicted * n_targets_predicted ) )\n\n    print()    \n    try:\n        print( Y.sum(),  Y.sum()/ (Y.shape[0] * Y.shape[1]) )\n    except:\n        pass    \n\n    print('Here is fast rationale why we should think of threshold for F1 is around 0.28 - number of 1 in that case corresponds to train data')\nexcept:\n    pass\n","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:19.608991Z","iopub.execute_input":"2023-08-03T09:48:19.609989Z","iopub.status.idle":"2023-08-03T09:48:19.61948Z","shell.execute_reply.started":"2023-08-03T09:48:19.609953Z","shell.execute_reply":"2023-08-03T09:48:19.617981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exercises. Take home messages. Check yourself questions","metadata":{}},{"cell_type":"markdown","source":"## Exercises","metadata":{}},{"cell_type":"markdown","source":"#### Ex 1 What is wrong with the following code ?  \n\n    for epoch in range(max_epoch):\n        \n        model.train() # switch model into train mode\n        for i_batch, (x_batch, y_batch) in enumerate(train_dataloader): # Loop ove batches\n            optimizer.zero_grad() # technical - set gradients to zero, otherwise they will be accumulated         \n            \n            preds = model(x_batch)# Compute predictions only for batch samples \n        \n            loss = criterion(preds, y_batch) # Compute loss function for batch predictions\n            \n            loss.backward() # Compute gradients\n            optimizer.step() # Update NN weights using gradients\n            \n            if lr_sched is not None:\n                lr_sched.step() # Step LR scheduler\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('%.1f seconds passed total '%(time.time()-t0start) )","metadata":{"execution":{"iopub.status.busy":"2023-08-03T09:48:19.62175Z","iopub.execute_input":"2023-08-03T09:48:19.622206Z","iopub.status.idle":"2023-08-03T09:48:19.631929Z","shell.execute_reply.started":"2023-08-03T09:48:19.622173Z","shell.execute_reply":"2023-08-03T09:48:19.630635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}