Applying Data Science For Social Good In Nonprofit Organization With Troubled Family Risk Profiling R Dashboard Application
Ting-Wei Lin1,Tsun-Wei Tseng2,Yu-Hsuan Lin3,Pei-Yu Chen4,Ning-Yuan Lyu5,Ting-Kuang Lo6,Shing-Yun Jung7,Hsu Wei8,Charles Chuang9,Chun-Yu Yin8,Johnson Hsieh10,11
- Genome and Systems Biology Degree Program, National Taiwan University and Academia Sinica
- TAO Info Co.Ltd
- inQ Technology Co.Ltd
- Pegatron Co.Ltd
- Department of Electrical Engineering, National Tsing-Hua University
- Department of applied mathematics, Feng Chia University
- Department of Computer Science, National Tsing-Hua University
- Hfoundataion
- NETivism Co.Ltd
- Department of Computer Science, National Chengchi University
- DSP Co.Ltd
Keywords: troubled family risk profiling, data science for social good, association rule analysis, steady-state analysis, topic model, shiny application
Webpages: https://weitinglin.shinyapps.io/d4sg_dashboard_v1/
Assessing the troubled families risk status and distributing the resources appropriately is a big issue for nonprofit organizations, not to mention the national program like UK Trouble Families Program^1^. Unfortunately, those organizations and national programs still use the conventional way to deal with these problems^2^. As the increasing demands for social assistance and longterm shortage of social workers in these field, a more precise and continuous way to handle the social resources wisely and efficiently is needed. Besides, the junior social workers are hard to quickly get a hang of several families through lots of interview records and past documents, then decide whether providing their intervention. And most importantly of all, they lack of ability to utilize those information with proper data engineering and summarize those experience through data analysis. So here We first apply a evidence-based approach on assessing the troubled families's risk status with a R dashboard application integrated the prediction model generated from those families archives, follow-up records under the Data Science For Social Good Program in Taiwan with the cooperation between a volunteer data science team and local nonprofit organization HFoundation.
In the beginning, the organization and volunteer data science team use customer journey analysis to map the social worker's experience and organization workflow in order to understand how organization generate their different types of data and define the proper scale of framework in dealing with their problems. Overall, There are now 57 families accepted active aid by organization and with maximum following time up to 6 years. The documents from these families have basic socioeconomic information, following interview records from home visits with various follow-up time by the social workers, mainly text files. After de-identification of those documents, we preprocess their family archives and cases follow-up interview into suitable format for further management. Then, We use risk factors system to tag each home visit interview events in order to create family risk prediction model. There are seven major risk factors categories range from financial problems to housing problems. And we also use topic model to extract more information related to those risk factors from families archives. Then, we use association rule analysis in those home visit records during these 6 years and discuss the result rules to the senior social workers and in the same time the steady state analysis with Markov chain were performed to calculate recurrence rate of high risk factors for each home visit event. Those result can help to detect possible underline risk factors and predict the possible recurrence rate for each risk factors from recent family status.
The dashboard shiny application was build with the above analysis result to assist the social worker on their daily work in recognizing the possible risk factors from those troubled families and prioritizing their resources to manage families high recurrence risk factors. The application provide a overview visualization of each case's family data in timeline with risk factors and basic summary statistic. Social workers can easily get the insights from cases families past records and know the possible underline risk factors with the association rules and decide which families's problems should tackle first with the highest recurrence rate predicted from the model. In addition, the social workers can input their home visit or interview data into the application which can update the model and become their day-to-day working tools and make their workflow more efficiently and precisely. In the end, the high-risk family profiling R dashboard application can be a great tool to provide those nonprofit organizations, not only HFoundation, in these field to precisely and efficiently manage their family cases and prioritizing their resources on managing families problems.
- Fletcher A, Gardner F, McKee M, Bonell C.(2012). The British government’s Troubled Families Programme. BMJ2012;344:e3403. doi: https://10.1136/bmj.e3403
- Chris Bonell, Martin McKee.(2016). Adam Fletcher. Troubled families, troubled policy making. BMJ 2016; 355 doi: https://doi.org/10.1136/bmj.i5879