GPT-4 and Code Interpreter are hotter than Barbie!

Happy Summer! I’m looking forward to seeing Oppenheimer and Barbie on vacation next week. Both are summer blockbusters which are hot, hot, hot! Explosive and blowing up everywhere. So… What could be bigger and a “real-life actual” game changer?
Generative Pre-trained Transformer (GPT-4) and Large Language Models (LLM’s).

It’s hard to describe how jaw-dropping Open AI GPT-4 Plus is and for only $20 per month how it can change your life. The ability to load a dataset, run analysis, plot the results, and have the python code available with narrative describing the rationale behind advances statistics is unbelievable. It’s clean, fast, and overall, technically accurate. Note: I’ve executed the python code generated by GPT-4 in my own Jupiter notebook and Visual Studio Code to check the results.

I’ll post additional thoughts on my node.js Azure sandbox but don’t wait for me – go get a subscription and try it out for yourself!

Data Warehouse, Data Lakehouse and Data Mesh

Last year, I read a very interesting blog post by Darwin Schweitzer a Microsoft technologist who discusses how to consider emerging technologies in the context of building sustainable enterprises. The blog posts relates the learning patterns organization adopt to newer data technologies replacing existing capability. The approach covers the strategic, organizational, architectural and technological challenges and changes with scaling enterprise analytics.

Three Horizon Model/Framework – strategic
Data Mesh Sociotechnical Paradigm – organizational
Data Lakehouse Architecture – architectural
Azure Cloud Scale Analytics Platform – technological

Read more by following this link:

https://techcommunity.microsoft.com/t5/data-architecture-blog/bring-vision-to-life-with-three-horizons-data-mesh-data/ba-p/3390414


https://geoparquet.org/ is worth checking out!

Recently, I watched a webinar hosted by TDWI, Databricks and Carto. The topic was Unlocking the Power of Spatial Analysis and Data Lakehouses. A copy of the webinar and the slide deck shared is available here. What I liked about the session was the use of Databricks and a Data Lake to provide Spatial Data. There was also a brief discussion on the role of the Open Geospatial Consortium. This group is working on the specifications for creating a geoparquet file. For anyone with an interest in GIS, Mapping, Data and Analytics this is worth checking out!

https://github.com/opengeospatial/geoparquet

ESRI Tapestry Segmentation

For the last 20 years, ESRI has been capturing geographic and demographic data. In the 90’s Acxiom (and others) came up with Lifestyle Segmentation. ESRI started in 1969 and today maps “everything” down to the household level.

I’m surprised by the accuracy of Tapestry Segmentation. The top three segments in my zip code are: “Top Tier”, “Comfortable Empty Nesters” and “In Style”. Pretty accurate – with one data point. I’m an Empty Nester who would like to be “In Style” or “Top Tier” but feel very lucky and fortunate to be “comfortable”.

The 45243 Zip Code

Here is a link to a PDF with more on Tapestry Segmentation and my personal segment.
Just in case you we’re interested here are the details on “Top Tier” and ‘In Style“.

What is Differential privacy?

Differential privacy seeks to protect individual data values by adding statistical “noise” to the analysis process. The math involved in adding the noise is complex, but the principle is fairly intuitive – the noise ensures that data aggregations stay statistically consistent with the actual data values allowing for some random variation, but make it impossible to work out the individual values from the aggregated data. In addition, the noise is different for each analysis, so the results are non-deterministic – in other words, two analyses that perform the same aggregation may produce slightly different results.

Understand Differential Privacy!

Microsoft Viva

Microsoft Viva is an employee experience platform that brings together communications, knowledge, learning, resources, and insights in the flow of work. Powered by Microsoft 365 and experienced through Microsoft Teams, Viva fosters a culture that empowers people and teams to be their best from anywhere.

Viva Learning
Viva Insights
Viva Topics
Viva Connections

https://www.thinkdataanalysis.com/microsoft_viva.html

Think Data Analysis.com

Recently I decided to move from AWS to Azure for the hosting of my “Sandbox” sites. With the move, I plan to add serve up live interactive data content highlighting different “data” projects of personal and professional interest.

Much of what I’ve worked on is contained on “corporate” portals and intranets. The move to Azure from AWS for “server based” content will allow more flexibility and access to Power BI, SharePoint and Microsoft Teams.